Claude Opus 5
Anthropic·Released Jul 24, 2026
Updated Oct 8, 11:55 PM ET
The score is on the board's scale (reference average 50, about 15 points per standard deviation), not a percentage.
By domain
Each domain is scored on the same scale as the overall score. Coverage is the share of the domain's kinds of evaluation with a result; it says how much evidence there is, not how good the model is.
Every result behind the score
The result as the evaluator publishes it, and the reading it gives on the board's scale. A hard evaluation can give a high reading for a modest result; results near 95% or 5% only say "at least" or "at most".
Coding
10 results · 74.1- APEX-SWEMercor · max effort63.7%Reading 86.8
- Vibe Code Bench 1–100Vals AI · max effort28.5%Reading 80.3
- Terminal-Bench 4.0Vals AI · max effort53.5%Reading 80.0
- Terminal-Bench 4.0 (AA run)Artificial Analysis · high effort46.0%Reading 78.6
- LiveBench CodingLiveBench · max effort73.3%Reading 77.0
- Code MigrationVals AI · max effort57.5%Reading 73.9
- Vibe Code BenchVals AI88.4%Reading 71.9
- CyberBench PatchVals AI · max effort85.7%Reading 68.5
- ProgramBenchVals AI · max effortNear the floor3.0%Reading at most 65.7
- SciCodeArtificial Analysis · high effort55.4%Reading 64.3
Research and reasoning
9 results · 75.6- ProofBenchVals AI · max effortNear the ceiling99.0%Reading at least 89.2
- Mystery Game PuzzlesEpoch AI · max effort59.0%Reading 81.1
- LiveBench ReasoningLiveBench · max effort93.5%Reading 77.5
- FrontierMath Tier 4Epoch AI · max effort73.2%Reading 75.0
- Terminal-Bench ScienceVals AI · max effort27.1%Reading 74.5
- FrontierMath Tiers 1–3Epoch AI · max effort85.6%Reading 73.9
- Humanity's Last Exam (AA run)Artificial Analysis · high effort52.8%Reading 73.5
- MysteryMechanismVals AI · max effort37.4%Reading 73.1
- Chess PuzzlesEpoch AI · max effort42.0%Reading 62.5
Professional work
9 results · 73.6- MedCodeVals AI · max effortCapped63.6%Reading 114.3
- Tax Agent BenchVals AI · max effort45.5%Reading 83.2
- Legal Research BenchVals AI · max effort55.3%Reading 78.9
- APEX-AgentsMercor · max effort65.8%Reading 76.4
- EMBVals AI · max effort73.6%Reading 73.8
- Finance AgentVals AI · max effort58.6%Reading 70.7
- τ-Bench Banking (AA run)Artificial Analysis · high effort44.7%Reading 68.9
- LiveBench Data AnalysisLiveBench · max effort74.6%Reading 56.2
- Harvey Legal Agent BenchmarkVals AI · max effort6.7%Reading 40.4
Knowledge and accuracy
5 results · 69.5- BullshitBenchBullshitBench · xhigh effort58.2%Reading 103.1
- LiveBench LanguageLiveBench · max effort88.7%Reading 88.5
- SimpleQA VerifiedEpoch AI · max effort59.9%Reading 68.8
- AA-LCRArtificial Analysis · high effort79.0%Reading 47.4
- LiveBench Instruction FollowingLiveBench · max effort63.8%Reading 37.9
Human preference
1 result · 59.2- Arena TextLMArena · high effort1,489.7Reading 52.3
Shown, not scored
Composite indices, saturated or older evaluations, evaluations still being checked, and vision, writing and multilingual results. They do not move the score.
- IOIVals AI · CodingReference only84.3%
- LiveCodeBenchVals AI · CodingReference only89.0%
- SkillsBenchVals AI · CodingWatching60.4%
- SRE BenchVals AI · CodingWatching12.2%
- SWE-bench VerifiedVals AI · CodingReference only97.0%
- BioMysteryBenchVals AI · Research and reasoningWatching79.3%
- OTIS Mock AIMEEpoch AI · Research and reasoningReference only98.9%
- APEX-AccountingMercor · Professional workWatching10.5%
- CorpFinVals AI · Professional workReference only73.2%
- EBR-benchEpoch AI · Professional workWatching45.7%
- LegalBenchVals AI · Professional workReference only87.0%
- MedScribeVals AI · Professional workWatching91.0%
- TaxEvalVals AI · Professional workReference only75.1%
- GPQA DiamondEpoch AI · Knowledge and accuracyReference only93.9%
- GPQA DiamondVals AI · Knowledge and accuracyReference only93.4%
- MMLU-ProVals AI · Knowledge and accuracyReference only91.6%
- AA Intelligence IndexArtificial Analysis · Composite indicesReference only48.1
- LiveBench averageLiveBench · Composite indicesReference only80.1%
- Vals IndexVals AI · Composite indicesReference only63.7%
- Arena VisionLMArena · VisionReference only1,289.0
- MMMU ProVals AI · VisionReference only89.9%
- Arena WebDevLMArena · Writing and designReference only1,657.7
No result yet
Scored evaluations this model has not taken. A missing result neither adds nor subtracts.
- τ²-Bench Telecom (AA run) · Artificial Analysis
- IFBench (AA run) · Artificial Analysis