Anthropic finds container resource limits can shift agentic coding eval scores
Original titleQuantifying infrastructure noise in agentic coding evals
AISummary
Anthropic reports that resource configuration alone can move Terminal-Bench 2.0 scores by up to 6 percentage points, with infra error rates falling from 5.8% under strict enforcement to 0.5% when uncapped. Above about 3x the per-task specs, extra headroom starts letting agents solve tasks they previously could not, so limits can change what the eval measures.
AIWhy it matters
The source shows how container resource limits shift agentic coding scores, which helps readers interpret small leaderboard gaps and set up evals more consistently.
Source: Anthropic Engineering · anthropic.comPublished · added here