Banger paper from Meta Superintelligence Labs on self-improving agents.
(bookmark it)
It's hard to know exactly what drives self-improvement, since so many variables are at play (notes, skills, tool calls between runs, etc.).
Meta researchers explore and discuss a way to measure whether self-improvement pays off.
They call it agent plasticity. It is the gain on held-out tasks per dollar spent on learning, with model weights frozen and every run starting from a fresh context.
They find that the model that performs the best is often a different model from the one that learns most efficiently.
In chess, Go, and Hex, Claude Fable 5 reaches the highest final score, while GPT-5.6 Sol gains the most per dollar.
In NetHack, only Claude Opus 5.5 improves significantly, by 66 normalized points for about $1,073 of learning.
Another interesting finding is that slow learners often ignore artifacts they already wrote. Faster learners reuse their artifacts and still fail when an artifact is low quality.
Paper: https://arxiv.org/abs/2610.08902
Chat with Paper: https://academy.dair.ai/papers/agent-plasticity-measuring-self-improvement-through-experience-2610.08902
