Skip to content
Read the original: Epoch AI· David Owen·Published PickAI score60/100

Epoch AI finds frontier models fall short of an end-to-end AI research task

Original titleCan AI automate AI R&D yet?

AISummary

Epoch AI's InnovationEval tested whether AI agents could independently devise a post-training method matching on-policy self-distillation (SDPO), a recent human-developed innovation.

GPT-5.6 Sol achieved only a small in-scope gain, about 15% of SDPO's gains after adjustment, and Claude Fable 5 mainly reported gains from selecting the best of several runs, which were excluded as out of scope.

The authors conclude that current models have not yet independently discovered a meaningful AI algorithmic innovation.

AIWhy it matters

The evaluation tests whether AI can independently devise a post-training method matching a published human innovation, with a scope and memorization caveat worth reading.

Read the original epoch.ai

Source: Epoch AI · epoch.ai