Skip to content
Read the original: OpenBMB· OpenBMB·Published· 10d agoPickAI score72

One-Shot OPD: One Training Query Matches Most of Full-Data Distillation Gains

Post-training pipelines now use on-policy distillation (OPD) to hand a student the teacher's full next-token distribution at every prefix...

AISummary

Researchers from Tsinghua NLP and collaborators show that on-policy distillation with a single training query recovers 87% of full-data gains on math, reaching 68.5 versus 69.8 by step 300. The paper attributes the slow progress to how fast the student absorbs the teacher's signal rather than to dataset size. Code and the paper are publicly available on GitHub and Hugging Face.

AIWhy it matters

The paper isolates training data from the algorithm, showing one query nearly matches full-data on-policy distillation, which reframes where post-training gains come from.

Read the original x.com

Source: OpenBMB · x.com