GPT-5.6 Sol
OpenAI·Released Jul 9, 2026
Updated Oct 8, 11:55 PM ET
The score is on the board's scale (reference average 50, about 15 points per standard deviation), not a percentage.
By domain
Each domain is scored on the same scale as the overall score. Coverage is the share of the domain's kinds of evaluation with a result; it says how much evidence there is, not how good the model is.
Every result behind the score
The result as the evaluator publishes it, and the reading it gives on the board's scale. A hard evaluation can give a high reading for a modest result; results near 95% or 5% only say "at least" or "at most".
Coding
10 results · 68.0- CyberBench PatchVals AI · max effort89.3%Reading 82.9
- SciCodeArtificial Analysis · high effort57.8%Reading 71.0
- Code MigrationVals AI · max effort52.9%Reading 70.3
- Terminal-Bench 4.0Vals AI · max effort37.9%Reading 70.3
- LiveBench CodingLiveBench · max effort70.1%Reading 67.6
- ProgramBenchVals AI · max effortNear the floor1.5%Reading at most 65.7
- Vibe Code Bench 1–100Vals AI · max effort20.0%Reading 65.4
- Vibe Code BenchVals AI · max effort80.5%Reading 64.3
- APEX-SWEMercor · xhigh effort45.8%Reading 60.8
- Terminal-Bench 4.0 (AA run)Artificial Analysis · high effort20.7%Reading 59.9
Research and reasoning
9 results · 74.0- FrontierMath Tier 4Epoch AI · max effort82.9%Reading 82.7
- Mystery Game PuzzlesEpoch AI · max effort58.0%Reading 80.1
- LiveBench ReasoningLiveBench · max effort93.9%Reading 79.9
- FrontierMath Tiers 1–3Epoch AI · max effort89.1%Reading 79.7
- Chess PuzzlesEpoch AI · max effort55.0%Reading 78.2
- ProofBenchVals AI · max effort83.0%Reading 73.9
- Terminal-Bench ScienceVals AI · max effort20.0%Reading 69.3
- MysteryMechanismVals AI · max effort33.3%Reading 68.8
- Humanity's Last Exam (AA run)Artificial Analysis · high effort46.0%Reading 64.7
Professional work
9 results · 58.5- LiveBench Data AnalysisLiveBench · max effort79.8%Reading 73.9
- EMBVals AI · max effort72.3%Reading 72.1
- Legal Research BenchVals AI · max effort48.1%Reading 71.0
- Tax Agent BenchVals AI · max effort31.1%Reading 64.6
- Finance AgentVals AI · max effort53.8%Reading 61.0
- τ-Bench Banking (AA run)Artificial Analysis · high effort36.7%Reading 58.6
- MedCodeVals AI · max effort44.0%Reading 47.6
- τ²-Bench Telecom (AA run)Artificial Analysis · high effort83.3%Reading 33.9
- Harvey Legal Agent BenchmarkVals AI · max effortCappedNear the floor2.5%Reading at most 24.8
Knowledge and accuracy
6 results · 60.1- LiveBench LanguageLiveBench · max effort87.7%Reading 84.1
- SimpleQA VerifiedEpoch AI · max effort69.7%Reading 84.0
- LiveBench Instruction FollowingLiveBench · max effort71.8%Reading 70.2
- AA-LCRArtificial Analysis · high effort81.7%Reading 52.0
- IFBench (AA run)Artificial Analysis · high effort69.2%Reading 39.5
- BullshitBenchBullshitBench · max effortCapped27.3%Reading 19.1
Human preference
1 result · 55.4- Arena TextLMArena · xhigh effort1,483.9Reading 50.6
Shown, not scored
Composite indices, saturated or older evaluations, evaluations still being checked, and vision, writing and multilingual results. They do not move the score.
- IOIVals AI · CodingReference only91.2%
- LiveCodeBenchVals AI · CodingReference only82.6%
- MirrorCodeEpoch AI · CodingWatching20.0%
- SkillsBenchVals AI · CodingWatching54.1%
- SRE BenchVals AI · CodingWatching30.5%
- SWE-bench VerifiedVals AI · CodingReference only96.2%
- BioMysteryBenchVals AI · Research and reasoningWatching71.1%
- OTIS Mock AIMEEpoch AI · Research and reasoningReference only100.0%
- CorpFinVals AI · Professional workReference only64.4%
- EBR-benchEpoch AI · Professional workWatching44.8%
- LegalBenchVals AI · Professional workReference only87.0%
- MedScribeVals AI · Professional workWatching85.2%
- TaxEvalVals AI · Professional workReference only74.8%
- GPQA DiamondEpoch AI · Knowledge and accuracyReference only93.5%
- GPQA DiamondVals AI · Knowledge and accuracyReference only95.2%
- MMLU-ProVals AI · Knowledge and accuracyReference only89.1%
- AA Intelligence IndexArtificial Analysis · Composite indicesReference only42.3
- LiveBench averageLiveBench · Composite indicesReference only81.1%
- Vals IndexVals AI · Composite indicesReference only58.0%
- Arena VisionLMArena · VisionReference only1,284.7
- MMMU ProVals AI · VisionReference only88.8%
No result yet
Scored evaluations this model has not taken. A missing result neither adds nor subtracts.
- APEX-Agents · Mercor