METR's preliminary evaluation finds Claude Opus 5.5 is an incremental AI R&D gain over Fable 5.1
Original titleSummary of METR's predeployment evaluation of Claude Opus 5.5
AISummary
METR's preliminary evaluation concludes that Claude Opus 5.5 likely gives slightly higher AI R&D productivity uplift than Fable 5.1 but is unlikely to fully automate AI R&D. The evaluation used five capability tasks over 10 business days of API access, and METR says Anthropic reviewed and edited the summary before sign-off.
AIWhy it matters
The report separates two claims about AI R&D acceleration and discloses that Anthropic reviewed the summary, which helps readers weigh its independence and evidence.
Source: METR Blog · metr.orgPublished · added here