Skip to content
Read the original: Anthropic Engineering· Published Pick72/100AI score72/100

Anthropic finds container resource limits can shift agentic coding eval scores

Original titleQuantifying infrastructure noise in agentic coding evals

AISummary

Anthropic reports that resource configuration alone can move Terminal-Bench 2.0 scores by up to 6 percentage points, with infra error rates falling from 5.8% under strict enforcement to 0.5% when uncapped. Above about 3x the per-task specs, extra headroom starts letting agents solve tasks they previously could not, so limits can change what the eval measures.

AIWhy it matters

The source shows how container resource limits shift agentic coding scores, which helps readers interpret small leaderboard gaps and set up evals more consistently.

Read the original anthropic.com

Source: Anthropic Engineering · anthropic.comPublished · added here