GPT-5.2
OpenAI·Released Dec 11, 2025
Updated Oct 8, 11:55 PM ET
The score is on the board's scale (reference average 50, about 15 points per standard deviation), not a percentage.
By domain
Each domain is scored on the same scale as the overall score. Coverage is the share of the domain's kinds of evaluation with a result; it says how much evidence there is, not how good the model is.
Every result behind the score
The result as the evaluator publishes it, and the reading it gives on the board's scale. A hard evaluation can give a high reading for a modest result; results near 95% or 5% only say "at least" or "at most".
Coding
2 results · 48.6- LiveBench CodingLiveBench · high effort63.2%Reading 49.3
- Vibe Code BenchVals AI · xhigh effort53.5%Reading 48.6
Research and reasoning
6 results · 52.6- Chess PuzzlesEpoch AI · high effort40.0%Reading 60.0
- LiveBench ReasoningLiveBench · high effort88.2%Reading 57.1
- FrontierMath Tiers 1–3Epoch AI · xhigh effort67.4%Reading 55.0
- Humanity's Last Exam (AA run)Artificial Analysis · xhigh effort37.7%Reading 53.5
- FrontierMath Tier 4Epoch AI · xhigh effort31.7%Reading 51.3
- Mystery Game PuzzlesEpoch AI · high effort23.0%Reading 42.2
Professional work
4 results · 52.0- LiveBench Data AnalysisLiveBench · high effort78.2%Reading 68.0
- MedCodeVals AI · xhigh effort49.7%Reading 67.0
- τ²-Bench Telecom (AA run)Artificial Analysis · xhigh effort84.8%Reading 36.0
- τ-Bench Banking (AA run)Artificial Analysis · none effort11.1%Reading 11.2
Knowledge and accuracy
6 results · 40.5- LiveBench LanguageLiveBench · high effort79.8%Reading 57.3
- AA-LCRArtificial Analysis · xhigh effort82.7%Reading 53.9
- IFBench (AA run)Artificial Analysis · xhigh effort75.4%Reading 53.6
- SimpleQA VerifiedEpoch AI · high effort34.3%Reading 31.8
- LiveBench Instruction FollowingLiveBench · high effort61.8%Reading 30.5
- BullshitBenchBullshitBench · high effortCapped23.6%Reading 6.9
Human preference
1 result · 40.1- Arena TextLMArena · high effort1,437.3Reading 36.7
Shown, not scored
Composite indices, saturated or older evaluations, evaluations still being checked, and vision, writing and multilingual results. They do not move the score.
- LiveCodeBenchVals AI · CodingReference only85.4%
- LiveCodeBench (AA run)Artificial Analysis · CodingReference only88.9%
- SWE-bench VerifiedEpoch AI · CodingReference only73.8%
- SWE-bench VerifiedVals AI · CodingReference only75.8%
- AIME (AA run)Artificial Analysis · Research and reasoningReference only99.0%
- OTIS Mock AIMEEpoch AI · Research and reasoningReference only96.1%
- CorpFinVals AI · Professional workReference only65.9%
- EBR-benchEpoch AI · Professional workWatching23.0%
- LegalBenchVals AI · Professional workReference only82.8%
- MedScribeVals AI · Professional workWatching84.4%
- TaxEvalVals AI · Professional workReference only75.8%
- GPQA DiamondEpoch AI · Knowledge and accuracyReference only88.2%
- GPQA DiamondVals AI · Knowledge and accuracyReference only91.7%
- MMLU-ProVals AI · Knowledge and accuracyReference only86.2%
- AA Intelligence IndexArtificial Analysis · Composite indicesReference only30.4
- LiveBench averageLiveBench · Composite indicesReference only74.6%
- Arena VisionLMArena · VisionReference only1,242.8
- MMMU ProVals AI · VisionReference only86.7%
- Arena WebDevLMArena · Writing and designReference only1,416.1
- MGSMVals AI · MultilingualReference only94.0%
No result yet
Scored evaluations this model has not taken. A missing result neither adds nor subtracts.
- Terminal-Bench 4.0 · Vals AI
- Terminal-Bench 4.0 (AA run) · Artificial Analysis
- Code Migration · Vals AI
- ProgramBench · Vals AI
- Vibe Code Bench 1–100 · Vals AI
- CyberBench Patch · Vals AI
- APEX-SWE · Mercor
- SciCode · Artificial Analysis
- MysteryMechanism · Vals AI
- Terminal-Bench Science · Vals AI
- ProofBench · Vals AI
- Finance Agent · Vals AI
- APEX-Agents · Mercor
- Harvey Legal Agent Benchmark · Vals AI
- Legal Research Bench · Vals AI
- Tax Agent Bench · Vals AI
- EMB · Vals AI