Skip to content
Read the original: Epoch AI· Published 22/100AI score22/100

Epoch AI notes benchmark gains from on-policy self-distillation (SDPO)

Original titleWe instructed the AI models that their technique should improve performance on several benchmarks. We already knew that all of these coul...

AISummary

Epoch AI reports that AI models instructed to improve on several benchmarks could already be boosted by a recent human-authored post-training method, on-policy self-distillation (SDPO). The post implies these gains were known before the models' own technique was evaluated.

Read the original x.com

Source: Epoch AI · x.comPublished · added here