Grok 4.6
xAI·Released Aug 12, 2026
Updated Oct 8, 11:55 PM ET
The score is on the board's scale (reference average 50, about 15 points per standard deviation), not a percentage.
By domain
Each domain is scored on the same scale as the overall score. Coverage is the share of the domain's kinds of evaluation with a result; it says how much evidence there is, not how good the model is.
Every result behind the score
The result as the evaluator publishes it, and the reading it gives on the board's scale. A hard evaluation can give a high reading for a modest result; results near 95% or 5% only say "at least" or "at most".
Coding
10 results · 61.3- APEX-SWEMercor · high effort56.4%Reading 76.0
- SciCodeArtificial Analysis · high effort56.5%Reading 67.3
- ProgramBenchVals AI · high effortNear the floor1.0%Reading at most 65.7
- Code MigrationVals AI · high effort44.6%Reading 64.0
- Vibe Code BenchVals AI · high effort76.2%Reading 61.2
- Terminal-Bench 4.0 (AA run)Artificial Analysis · high effort21.2%Reading 60.4
- LiveBench CodingLiveBench66.9%Reading 59.0
- Terminal-Bench 4.0Vals AI · high effort17.2%Reading 53.9
- Vibe Code Bench 1–100Vals AI · high effort14.8%Reading 53.6
- CyberBench PatchVals AI · high effort80.4%Reading 51.6
Research and reasoning
9 results · 60.2- LiveBench ReasoningLiveBench91.5%Reading 68.7
- MysteryMechanismVals AI · high effort30.6%Reading 65.7
- Terminal-Bench ScienceVals AI · high effort11.4%Reading 60.7
- Humanity's Last Exam (AA run)Artificial Analysis · high effort42.9%Reading 60.6
- Chess PuzzlesEpoch AI · high effort40.0%Reading 60.0
- ProofBenchVals AI · high effort51.0%Reading 56.5
- Mystery Game PuzzlesEpoch AI · xhigh effort34.0%Reading 55.7
- FrontierMath Tiers 1–3Epoch AI · xhigh effort66.0%Reading 53.8
- FrontierMath Tier 4Epoch AI · xhigh effort31.7%Reading 51.3
Professional work
9 results · 67.1- Harvey Legal Agent BenchmarkVals AI · high effort15.8%Reading 90.0
- τ-Bench Banking (AA run)Artificial Analysis · high effort50.7%Reading 76.4
- APEX-AgentsMercor · xhigh effort65.3%Reading 75.8
- Legal Research BenchVals AI · high effort48.1%Reading 71.0
- Tax Agent BenchVals AI · high effort33.1%Reading 67.5
- Finance AgentVals AI · high effort53.7%Reading 60.9
- EMBVals AI · high effort62.7%Reading 60.0
- LiveBench Data AnalysisLiveBench73.9%Reading 54.1
- MedCodeVals AI · high effort44.7%Reading 50.1
Knowledge and accuracy
5 results · 68.3- BullshitBenchBullshitBench · max effortCapped60.0%Reading 107.9
- LiveBench Instruction FollowingLiveBench71.9%Reading 70.3
- LiveBench LanguageLiveBench83.7%Reading 69.2
- SimpleQA VerifiedEpoch AI · high effort49.3%Reading 53.7
- AA-LCRArtificial Analysis · high effort80.3%Reading 49.7
Human preference
1 result · 48.8- Arena TextLMArena · high effort1,453.8Reading 41.6
Shown, not scored
Composite indices, saturated or older evaluations, evaluations still being checked, and vision, writing and multilingual results. They do not move the score.
- IOIVals AI · CodingReference only47.6%
- LiveCodeBenchVals AI · CodingReference only88.2%
- SkillsBenchVals AI · CodingWatching55.8%
- SWE-bench VerifiedVals AI · CodingReference only95.6%
- BioMysteryBenchVals AI · Research and reasoningWatching72.2%
- OTIS Mock AIMEEpoch AI · Research and reasoningReference only97.8%
- CorpFinVals AI · Professional workReference only66.2%
- EBR-benchEpoch AI · Professional workWatching30.5%
- LegalBenchVals AI · Professional workReference only86.3%
- MedScribeVals AI · Professional workWatching86.5%
- TaxEvalVals AI · Professional workReference only71.1%
- GPQA DiamondEpoch AI · Knowledge and accuracyReference only94.0%
- GPQA DiamondVals AI · Knowledge and accuracyReference only94.7%
- MMLU-ProVals AI · Knowledge and accuracyReference only89.4%
- AA Intelligence IndexArtificial Analysis · Composite indicesReference only44.3
- LiveBench averageLiveBench · Composite indicesReference only78.0%
- Vals IndexVals AI · Composite indicesReference only52.1%
- Arena VisionLMArena · VisionReference only1,263.5
- Arena WebDevLMArena · Writing and designReference only1,617.0
No result yet
Scored evaluations this model has not taken. A missing result neither adds nor subtracts.
- τ²-Bench Telecom (AA run) · Artificial Analysis
- IFBench (AA run) · Artificial Analysis