Epoch AI finds frontier models fall short of an end-to-end AI research task
Original titleCan AI automate AI R&D yet?
AISummary
Epoch AI's InnovationEval tested whether AI agents could independently devise a post-training method matching on-policy self-distillation (SDPO), a recent human-developed innovation.
GPT-5.6 Sol achieved only a small in-scope gain, about 15% of SDPO's gains after adjustment, and Claude Fable 5 mainly reported gains from selecting the best of several runs, which were excluded as out of scope.
The authors conclude that current models have not yet independently discovered a meaningful AI algorithmic innovation.
AIWhy it matters
The evaluation tests whether AI can independently devise a post-training method matching a published human innovation, with a scope and memorization caveat worth reading.
Source: Epoch AI · epoch.ai