Claude Sonnet 5
Anthropic·Released Jun 30, 2026
Updated Oct 8, 11:55 PM ET
The score is on the board's scale (reference average 50, about 15 points per standard deviation), not a percentage.
By domain
Each domain is scored on the same scale as the overall score. Coverage is the share of the domain's kinds of evaluation with a result; it says how much evidence there is, not how good the model is.
Every result behind the score
The result as the evaluator publishes it, and the reading it gives on the board's scale. A hard evaluation can give a high reading for a modest result; results near 95% or 5% only say "at least" or "at most".
Coding
10 results · 58.3- LiveBench CodingLiveBench · xhigh effort70.0%Reading 67.5
- ProgramBenchVals AI · max effortNear the floor0.0%Reading at most 65.7
- Vibe Code BenchVals AI · max effort81.3%Reading 65.0
- Code MigrationVals AI · max effort44.4%Reading 63.8
- APEX-SWEMercor · max effort46.4%Reading 61.7
- SciCodeArtificial Analysis · high effort54.3%Reading 61.3
- CyberBench PatchVals AI · max effort82.1%Reading 56.8
- Vibe Code Bench 1–100Vals AI · max effort13.8%Reading 51.2
- Terminal-Bench 4.0Vals AI · max effort9.6%Reading 43.7
- Terminal-Bench 4.0 (AA run)Artificial Analysis · high effort5.1%Reading 34.7
Research and reasoning
8 results · 56.9- ProofBenchVals AI · max effort77.0%Reading 69.7
- LiveBench ReasoningLiveBench · xhigh effort90.8%Reading 65.9
- Mystery Game PuzzlesEpoch AI · max effort35.0%Reading 56.8
- Chess PuzzlesEpoch AI · xhigh effort35.0%Reading 53.6
- FrontierMath Tiers 1–3Epoch AI · max effort65.6%Reading 53.5
- Terminal-Bench ScienceVals AI · max effort5.7%Reading 50.8
- Humanity's Last Exam (AA run)Artificial Analysis · high effort35.7%Reading 50.7
- FrontierMath Tier 4Epoch AI · max effort29.3%Reading 49.8
Professional work
9 results · 56.2- EMBVals AI · max effort66.3%Reading 64.3
- Legal Research BenchVals AI · max effort41.8%Reading 64.1
- APEX-AgentsMercor · max effort54.5%Reading 62.3
- Tax Agent BenchVals AI · max effort29.4%Reading 62.2
- Finance AgentVals AI · max effort53.9%Reading 61.3
- MedCodeVals AI · max effort47.5%Reading 59.6
- τ-Bench Banking (AA run)Artificial Analysis · max effort37.3%Reading 59.4
- LiveBench Data AnalysisLiveBench · xhigh effort71.7%Reading 47.8
- Harvey Legal Agent BenchmarkVals AI · max effortNear the floor5.0%Reading at most 24.8
Knowledge and accuracy
5 results · 51.4- BullshitBenchBullshitBench · max effortCapped61.8%Reading 112.8
- LiveBench LanguageLiveBench · xhigh effort75.0%Reading 44.7
- AA-LCRArtificial Analysis · high effort76.7%Reading 43.8
- LiveBench Instruction FollowingLiveBench · xhigh effort63.9%Reading 38.3
- SimpleQA VerifiedEpoch AI · xhigh effort32.9%Reading 29.6
Human preference
1 result · 47.9- Arena TextLMArena · high effort1,461.8Reading 44.0
Shown, not scored
Composite indices, saturated or older evaluations, evaluations still being checked, and vision, writing and multilingual results. They do not move the score.
- IOIVals AI · CodingReference only45.0%
- LiveCodeBenchVals AI · CodingReference only82.4%
- SkillsBenchVals AI · CodingWatching46.5%
- SRE BenchVals AI · CodingWatching0.4%
- SWE-bench VerifiedVals AI · CodingReference only79.6%
- OTIS Mock AIMEEpoch AI · Research and reasoningReference only94.7%
- CorpFinVals AI · Professional workReference only67.0%
- LegalBenchVals AI · Professional workReference only83.9%
- MedScribeVals AI · Professional workWatching76.1%
- TaxEvalVals AI · Professional workReference only75.6%
- GPQA DiamondEpoch AI · Knowledge and accuracyReference only90.5%
- GPQA DiamondVals AI · Knowledge and accuracyReference only88.9%
- MMLU-ProVals AI · Knowledge and accuracyReference only87.5%
- AA Intelligence IndexArtificial Analysis · Composite indicesReference only31.7
- LiveBench averageLiveBench · Composite indicesReference only76.0%
- Vals IndexVals AI · Composite indicesReference only51.8%
- Arena VisionLMArena · VisionReference only1,264.5
- MMMU ProVals AI · VisionReference only83.0%
- Arena WebDevLMArena · Writing and designReference only1,540.9
No result yet
Scored evaluations this model has not taken. A missing result neither adds nor subtracts.
- MysteryMechanism · Vals AI
- τ²-Bench Telecom (AA run) · Artificial Analysis
- IFBench (AA run) · Artificial Analysis