Claude Fable 5.1
Anthropic·Released Sep 1, 2026
Updated Oct 8, 11:55 PM ET
The score is on the board's scale (reference average 50, about 15 points per standard deviation), not a percentage.
By domain
Each domain is scored on the same scale as the overall score. Coverage is the share of the domain's kinds of evaluation with a result; it says how much evidence there is, not how good the model is.
Every result behind the score
The result as the evaluator publishes it, and the reading it gives on the board's scale. A hard evaluation can give a high reading for a modest result; results near 95% or 5% only say "at least" or "at most".
Coding
10 results · 78.4- APEX-SWEMercor · max effort63.6%Reading 86.6
- LiveBench CodingLiveBench · max effort76.2%Reading 86.1
- Terminal-Bench 4.0Vals AI · max effort58.1%Reading 82.8
- Terminal-Bench 4.0 (AA run)Artificial Analysis · high effort52.0%Reading 82.4
- Vibe Code Bench 1–100Vals AI · max effort28.0%Reading 79.4
- ProgramBenchVals AI · max effort7.0%Reading 76.5
- CyberBench PatchVals AI · max effort87.5%Reading 75.2
- Vibe Code BenchVals AI · max effort90.3%Reading 74.3
- SciCodeArtificial Analysis · high effort58.7%Reading 73.5
- Code MigrationVals AI · max effort54.6%Reading 71.7
Research and reasoning
9 results · 80.8- ProofBenchVals AI · max effortNear the ceiling100.0%Reading at least 89.2
- FrontierMath Tier 4Epoch AI · max effort87.8%Reading 88.0
- MysteryMechanismVals AI · max effort47.7%Reading 83.5
- LiveBench ReasoningLiveBench · max effort94.3%Reading 82.3
- Terminal-Bench ScienceVals AI · max effort40.0%Reading 82.2
- FrontierMath Tiers 1–3Epoch AI · max effort90.2%Reading 81.7
- Mystery Game PuzzlesEpoch AI · max effort58.0%Reading 80.1
- Humanity's Last Exam (AA run)Artificial Analysis · high effort55.9%Reading 77.6
- Chess PuzzlesEpoch AI · max effort47.0%Reading 68.6
Professional work
9 results · 72.2- Tax Agent BenchVals AI · max effort49.2%Reading 87.6
- MedCodeVals AI · max effort53.5%Reading 79.5
- Legal Research BenchVals AI · max effort55.3%Reading 78.9
- EMBVals AI · max effort76.7%Reading 78.4
- LiveBench Data AnalysisLiveBench · max effort80.3%Reading 75.5
- Finance AgentVals AI · max effort58.9%Reading 71.2
- APEX-AgentsMercor · high effort59.7%Reading 68.7
- τ-Bench Banking (AA run)Artificial Analysis · high effort43.1%Reading 66.9
- Harvey Legal Agent BenchmarkVals AI · max effort6.7%Reading 40.4
Knowledge and accuracy
5 results · 82.9- BullshitBenchBullshitBench · max effort60.0%Reading 107.9
- LiveBench LanguageLiveBench · max effort89.5%Reading 92.3
- SimpleQA VerifiedEpoch AI · max effort70.8%Reading 85.8
- LiveBench Instruction FollowingLiveBench · max effort73.0%Reading 75.2
- AA-LCRArtificial Analysis · high effort83.7%Reading 55.8
Human preference
1 result · 62.9- Arena TextLMArena · max effort1,501.1Reading 55.7
Shown, not scored
Composite indices, saturated or older evaluations, evaluations still being checked, and vision, writing and multilingual results. They do not move the score.
- IOIVals AI · CodingReference only90.8%
- LiveCodeBenchVals AI · CodingReference only90.5%
- MirrorCodeEpoch AI · CodingWatching73.3%
- SkillsBenchVals AI · CodingWatching61.6%
- SRE BenchVals AI · CodingWatching22.9%
- OTIS Mock AIMEEpoch AI · Research and reasoningReference only100.0%
- APEX-AccountingMercor · Professional workWatching11.7%
- EBR-benchEpoch AI · Professional workWatching57.1%
- LegalBenchVals AI · Professional workReference only88.5%
- MedScribeVals AI · Professional workWatching91.3%
- TaxEvalVals AI · Professional workReference only76.0%
- GPQA DiamondVals AI · Knowledge and accuracyReference only93.4%
- MMLU-ProVals AI · Knowledge and accuracyReference only92.4%
- AA Intelligence IndexArtificial Analysis · Composite indicesReference only51.2
- LiveBench averageLiveBench · Composite indicesReference only83.4%
- Vals IndexVals AI · Composite indicesReference only65.8%
- Arena VisionLMArena · VisionReference only1,287.7
- MMMU ProVals AI · VisionReference only90.6%
- Arena WebDevLMArena · Writing and designReference only1,744.8
No result yet
Scored evaluations this model has not taken. A missing result neither adds nor subtracts.
- τ²-Bench Telecom (AA run) · Artificial Analysis
- IFBench (AA run) · Artificial Analysis