Grok 4.5
xAI·Released Jul 8, 2026
Updated Oct 8, 9:18 PM ET
The score is on the board's scale (reference average 50, about 15 points per standard deviation), not a percentage.
By domain
Each domain is scored on the same scale as the overall score. Coverage is the share of the domain's kinds of evaluation with a result; it says how much evidence there is, not how good the model is.
Every result behind the score
The result as the evaluator publishes it, and the reading it gives on the board's scale. A hard evaluation can give a high reading for a modest result; results near 95% or 5% only say "at least" or "at most".
Coding
9 results · 56.2- APEX-SWEMercor · high effort53.6%Reading 71.9
- ProgramBenchVals AI · high effortNear the floor0.0%Reading at most 65.7
- SciCodeArtificial Analysis · high effort55.0%Reading 63.2
- Code MigrationVals AI · high effort36.6%Reading 57.6
- Vibe Code BenchVals AI · high effort69.0%Reading 56.7
- CyberBench PatchVals AI · high effort80.4%Reading 51.6
- LiveBench CodingLiveBench62.5%Reading 47.7
- Terminal-Bench 4.0 (AA run)Artificial Analysis · high effort10.6%Reading 47.4
- Terminal-Bench 4.0Vals AI · high effort8.6%Reading 41.9
Research and reasoning
6 results · 54.1- Humanity's Last Exam (AA run)Artificial Analysis · high effort42.7%Reading 60.3
- LiveBench ReasoningLiveBench89.0%Reading 59.6
- Chess PuzzlesEpoch AI · high effort36.0%Reading 54.9
- FrontierMath Tiers 1–3Epoch AI · high effort57.2%Reading 47.1
- ProofBenchVals AI · high effort31.0%Reading 47.1
- FrontierMath Tier 4Epoch AI · high effort24.4%Reading 46.4
Professional work
9 results · 57.7- Harvey Legal Agent BenchmarkVals AI · high effort12.9%Reading 77.8
- τ-Bench Banking (AA run)Artificial Analysis · high effort42.1%Reading 65.6
- APEX-AgentsMercor · high effort56.2%Reading 64.4
- Legal Research BenchVals AI · high effort38.0%Reading 59.7
- Tax Agent BenchVals AI · high effort24.9%Reading 55.3
- LiveBench Data AnalysisLiveBench73.0%Reading 51.6
- Finance AgentVals AI · high effort48.3%Reading 50.5
- EMBVals AI · high effort52.9%Reading 48.9
- MedCodeVals AI · high effort43.3%Reading 45.3
Knowledge and accuracy
5 results · 65.5- BullshitBenchBullshitBench · high effortCapped60.0%Reading 107.9
- LiveBench Instruction FollowingLiveBench71.5%Reading 68.9
- LiveBench LanguageLiveBench82.8%Reading 66.3
- SimpleQA VerifiedEpoch AI · high effort48.3%Reading 52.3
- AA-LCRArtificial Analysis · high effort79.3%Reading 48.0
Human preference
1 result · 49.3- Arena TextLMArena1,465.8Reading 45.2
Shown, not scored
Composite indices, saturated or older evaluations, evaluations still being checked, and vision, writing and multilingual results. They do not move the score.
- IOIVals AI · CodingReference only40.6%
- LiveCodeBenchVals AI · CodingReference only87.4%
- SkillsBenchVals AI · CodingWatching66.0%
- SRE BenchVals AI · CodingWatching0.8%
- SWE-bench VerifiedVals AI · CodingReference only86.6%
- OTIS Mock AIMEEpoch AI · Research and reasoningReference only97.8%
- APEX-AccountingMercor · Professional workWatching4.5%
- CorpFinVals AI · Professional workReference only67.4%
- LegalBenchVals AI · Professional workReference only86.0%
- MedScribeVals AI · Professional workWatching86.9%
- TaxEvalVals AI · Professional workReference only71.7%
- GPQA DiamondEpoch AI · Knowledge and accuracyReference only93.4%
- GPQA DiamondVals AI · Knowledge and accuracyReference only92.9%
- MMLU-ProVals AI · Knowledge and accuracyReference only89.2%
- AA Intelligence IndexArtificial Analysis · Composite indicesReference only38.8
- LiveBench averageLiveBench · Composite indicesReference only75.8%
- Vals IndexVals AI · Composite indicesReference only44.7%
- Arena VisionLMArena · VisionReference only1,279.0
- MMMU ProVals AI · VisionReference only61.8%
- Arena WebDevLMArena · Writing and designReference only1,552.9
No result yet
Scored evaluations this model has not taken. A missing result neither adds nor subtracts.
- Vibe Code Bench 1–100 · Vals AI
- Mystery Game Puzzles · Epoch AI
- MysteryMechanism · Vals AI
- Terminal-Bench Science · Vals AI
- τ²-Bench Telecom (AA run) · Artificial Analysis
- IFBench (AA run) · Artificial Analysis