On-Policy Distillation Is Data-Overfed, Algorithm-Starved
New research reveals that on-policy distillation (OPD) for LLMs can achieve near-full performance using only a single training query. The findings suggest that current training pipelines are bottlenecked by slow student learning rather than a lack of diverse data.