Claude Fable 5
Anthropic·Released Jun 9, 2026
Updated Oct 8, 9:18 PM ET
The score is on the board's scale (reference average 50, about 15 points per standard deviation), not a percentage.
By domain
Each domain is scored on the same scale as the overall score. Coverage is the share of the domain's kinds of evaluation with a result; it says how much evidence there is, not how good the model is.
Every result behind the score
The result as the evaluator publishes it, and the reading it gives on the board's scale. A hard evaluation can give a high reading for a modest result; results near 95% or 5% only say "at least" or "at most".
Coding
8 results · 74.9- SciCodeArtificial Analysis · max effort61.0%Reading 80.0
- APEX-SWEMercor · max effort58.8%Reading 79.5
- LiveBench CodingLiveBench · max effort74.1%Reading 79.3
- Terminal-Bench 4.0 (AA run)Artificial Analysis · max effort42.4%Reading 76.3
- Vibe Code BenchVals AI · max effort90.4%Reading 74.4
- Terminal-Bench 4.0Vals AI · max effort41.4%Reading 72.6
- Code MigrationVals AI · max effort55.1%Reading 72.0
- ProgramBenchVals AI · max effortNear the floor2.0%Reading at most 65.7
Research and reasoning
8 results · 74.9- FrontierMath Tier 4Epoch AI · max effort90.2%Reading 91.3
- ProofBenchVals AI · max effortNear the ceiling95.0%Reading at least 89.2
- Humanity's Last Exam (AA run)Artificial Analysis · max effort55.5%Reading 77.1
- FrontierMath Tiers 1–3Epoch AI · max effort87.0%Reading 76.0
- LiveBench ReasoningLiveBench · max effort92.8%Reading 74.3
- Mystery Game PuzzlesEpoch AI · max effort52.0%Reading 74.0
- Terminal-Bench ScienceVals AI · max effort15.7%Reading 65.5
- Chess PuzzlesEpoch AI · high effort41.0%Reading 61.2
Professional work
10 results · 73.6- MedCodeVals AI · max effort56.1%Reading 88.2
- LiveBench Data AnalysisLiveBench · max effort80.5%Reading 76.5
- Tax Agent BenchVals AI · max effort38.3%Reading 74.2
- EMBVals AI · max effort73.7%Reading 74.0
- APEX-AgentsMercor · max effort63.6%Reading 73.5
- Legal Research BenchVals AI · max effort49.5%Reading 72.6
- Harvey Legal Agent BenchmarkVals AI · max effort11.3%Reading 69.7
- Finance AgentVals AI · max effort56.3%Reading 66.0
- τ-Bench Banking (AA run)Artificial Analysis · max effort38.1%Reading 60.5
- τ²-Bench Telecom (AA run)Artificial Analysis · max effortNear the ceiling98.5%Reading at least 60.0
Knowledge and accuracy
5 results · 71.9- LiveBench LanguageLiveBench · max effort90.7%Reading 98.3
- LiveBench Instruction FollowingLiveBench · max effort75.8%Reading 87.9
- SimpleQA VerifiedEpoch AI · xhigh effort70.7%Reading 85.7
- AA-LCRArtificial Analysis · max effort82.3%Reading 53.2
- IFBench (AA run)Artificial Analysis · max effortCapped63.5%Reading 28.0
Human preference
1 result · 62.2- Arena TextLMArena · high effort1,504.3Reading 56.6
Shown, not scored
Composite indices, saturated or older evaluations, evaluations still being checked, and vision, writing and multilingual results. They do not move the score.
- LiveCodeBenchVals AI · CodingReference only89.8%
- MirrorCodeEpoch AI · CodingWatching63.9%
- SWE-bench VerifiedVals AI · CodingReference only95.0%
- OTIS Mock AIMEEpoch AI · Research and reasoningReference only100.0%
- APEX-AccountingMercor · Professional workWatching9.8%
- CorpFinVals AI · Professional workReference only71.8%
- EBR-benchEpoch AI · Professional workWatching39.5%
- LegalBenchVals AI · Professional workReference only88.6%
- MedScribeVals AI · Professional workWatching88.5%
- TaxEvalVals AI · Professional workReference only76.9%
- GPQA DiamondEpoch AI · Knowledge and accuracyReference only83.3%
- GPQA DiamondVals AI · Knowledge and accuracyReference only93.2%
- MMLU-ProVals AI · Knowledge and accuracyReference only91.5%
- AA Intelligence IndexArtificial Analysis · Composite indicesReference only49.6
- LiveBench averageLiveBench · Composite indicesReference only83.0%
- Vals IndexVals AI · Composite indicesReference only61.4%
- Arena VisionLMArena · VisionReference only1,308.8
- MMMU ProVals AI · VisionReference only89.3%
- Arena WebDevLMArena · Writing and designReference only1,624.1
No result yet
Scored evaluations this model has not taken. A missing result neither adds nor subtracts.
- Vibe Code Bench 1–100 · Vals AI
- CyberBench Patch · Vals AI
- MysteryMechanism · Vals AI
- BullshitBench · BullshitBench