Epoch AI notes benchmark gains from on-policy self-distillation (SDPO)
Original titleWe instructed the AI models that their technique should improve performance on several benchmarks. We already knew that all of these coul...
AISummary
Epoch AI reports that AI models instructed to improve on several benchmarks could already be boosted by a recent human-authored post-training method, on-policy self-distillation (SDPO). The post implies these gains were known before the models' own technique was evaluated.
Source: Epoch AI · x.comPublished · added here