Claude Opus 4.8
Anthropic·Released May 28, 2026
Updated Oct 8, 11:55 PM ET
The score is on the board's scale (reference average 50, about 15 points per standard deviation), not a percentage.
By domain
Each domain is scored on the same scale as the overall score. Coverage is the share of the domain's kinds of evaluation with a result; it says how much evidence there is, not how good the model is.
Every result behind the score
The result as the evaluator publishes it, and the reading it gives on the board's scale. A hard evaluation can give a high reading for a modest result; results near 95% or 5% only say "at least" or "at most".
Coding
8 results · 61.3- Vibe Code BenchVals AI · max effort82.7%Reading 66.2
- Code MigrationVals AI · max effort47.3%Reading 66.0
- ProgramBenchVals AI · max effortNear the floor1.0%Reading at most 65.7
- SciCodeArtificial Analysis · max effort54.4%Reading 61.6
- Terminal-Bench 4.0 (AA run)Artificial Analysis · max effort21.7%Reading 60.8
- Terminal-Bench 4.0Vals AI · max effort23.2%Reading 59.7
- APEX-SWEMercor · max effort43.9%Reading 58.1
- LiveBench CodingLiveBench · max effort66.2%Reading 57.0
Research and reasoning
7 results · 59.5- LiveBench ReasoningLiveBench · max effort91.8%Reading 69.6
- Humanity's Last Exam (AA run)Artificial Analysis · max effort48.7%Reading 68.2
- FrontierMath Tiers 1–3Epoch AI · max effort80.0%Reading 66.8
- FrontierMath Tier 4Epoch AI · max effort56.1%Reading 64.9
- Mystery Game PuzzlesEpoch AI · xhigh effort31.0%Reading 52.3
- Chess PuzzlesEpoch AI · max effort34.0%Reading 52.2
- Terminal-Bench ScienceVals AI · max effortNear the floor4.3%Reading at most 49.0
Professional work
10 results · 60.2- MedCodeVals AI · max effort53.2%Reading 78.6
- EMBVals AI · max effort69.4%Reading 68.2
- Legal Research BenchVals AI · max effort43.8%Reading 66.2
- Tax Agent BenchVals AI · max effort30.0%Reading 63.1
- Finance AgentVals AI · max effort53.9%Reading 61.3
- Harvey Legal Agent BenchmarkVals AI · max effort9.6%Reading 60.6
- τ²-Bench Telecom (AA run)Artificial Analysis · max effort94.4%Reading 57.8
- APEX-AgentsMercor · max effort48.9%Reading 55.7
- τ-Bench Banking (AA run)Artificial Analysis · max effort34.2%Reading 55.3
- LiveBench Data AnalysisLiveBench · max effort66.0%Reading 32.2
Knowledge and accuracy
6 results · 59.8- BullshitBenchBullshitBench · xhigh effortCapped92.7%Reading 245.0
- LiveBench Instruction FollowingLiveBench · max effort72.0%Reading 71.0
- SimpleQA VerifiedEpoch AI · max effort53.0%Reading 58.9
- LiveBench LanguageLiveBench · max effort79.7%Reading 56.9
- AA-LCRArtificial Analysis · max effort77.7%Reading 45.3
- IFBench (AA run)Artificial Analysis · max effort62.2%Reading 25.6
Human preference
1 result · 53.2- Arena TextLMArena · high effort1,481.5Reading 49.9
Shown, not scored
Composite indices, saturated or older evaluations, evaluations still being checked, and vision, writing and multilingual results. They do not move the score.
- LiveCodeBenchVals AI · CodingReference only87.8%
- SkillsBenchVals AI · CodingWatching59.2%
- SWE-bench VerifiedVals AI · CodingReference only88.6%
- OTIS Mock AIMEEpoch AI · Research and reasoningReference only98.3%
- APEX-AccountingMercor · Professional workWatching4.2%
- CorpFinVals AI · Professional workReference only66.7%
- EBR-benchEpoch AI · Professional workWatching28.6%
- LegalBenchVals AI · Professional workReference only83.6%
- MedScribeVals AI · Professional workWatching85.8%
- TaxEvalVals AI · Professional workReference only75.6%
- GPQA DiamondEpoch AI · Knowledge and accuracyReference only91.0%
- GPQA DiamondVals AI · Knowledge and accuracyReference only92.4%
- MMLU-ProVals AI · Knowledge and accuracyReference only89.6%
- AA Intelligence IndexArtificial Analysis · Composite indicesReference only41.8
- LiveBench averageLiveBench · Composite indicesReference only76.2%
- Vals IndexVals AI · Composite indicesReference only55.1%
- Arena VisionLMArena · VisionReference only1,286.4
- MMMU ProVals AI · VisionReference only86.6%
- Arena WebDevLMArena · Writing and designReference only1,556.1
No result yet
Scored evaluations this model has not taken. A missing result neither adds nor subtracts.
- Vibe Code Bench 1–100 · Vals AI
- CyberBench Patch · Vals AI
- MysteryMechanism · Vals AI
- ProofBench · Vals AI