Skip to content
View original post on X: Lilian Weng· 44/100AI score44/100

On-policy distillation uses a teacher model as dense process reward

AISummary

Lilian Weng says on-policy distillation lets a teacher model act as a process reward model, providing dense rewards during training. The approach also prevents the out-of-distribution shock that SFT-style training can cause during rollouts. Thinking Machines' related post reports it outperforms other approaches for math reasoning and an internal chat assistant at a fraction of the cost.

Post on XView on X
@lilianweng

On-policy distillation provides an elegant way to use the teacher model as a process reward model to provide dense reward while preventing SFT style "OOD shock" during rollout.

Thinking Machines@thinkymachines
Our latest post explores on-policy distillation, a training approach that unites the error-correcting relevance of RL with the reward density of SFT. When training it for math reasoning and as an internal chat assistant, we find that on-policy distillation can outperform other approaches for a fraction of the cost. https://thinkingmachines.ai/blog/on-policy-distillation/
View quoted post on X

Source: Lilian Weng · x.comPublished · added here