GPT-6.1 Sol
OpenAI·Released Sep 29, 2026
Updated Oct 8, 21:18 ET
The score is on the board's scale (reference average 50, about 15 points per standard deviation), not a percentage.
By domain
Each domain is scored on the same scale as the overall score. Coverage is the share of the domain's kinds of evaluation with a result; it says how much evidence there is, not how good the model is.
Every result behind the score
The result as the evaluator publishes it, and the reading it gives on the board's scale. A hard evaluation can give a high reading for a modest result; results near 95% or 5% only say "at least" or "at most".
Coding
9 results · 67.5- Terminal-Bench 4.0 (AA run)Artificial Analysis · high effort51.5%Reading 82.1
- Terminal-Bench 4.0Vals AI · max effort55.1%Reading 81.0
- Code MigrationVals AI · max effort65.1%Reading 80.1
- Vibe Code BenchVals AI · max effort88.9%Reading 72.5
- ProgramBenchVals AI · max effortNear the floor3.0%Reading at most 65.7
- SciCodeArtificial Analysis · high effort55.8%Reading 65.4
- LiveBench CodingLiveBench · xhigh effort68.8%Reading 64.0
- APEX-SWEMercor · max effort46.9%Reading 62.4
- CyberBench PatchVals AI · max effort78.6%Reading 46.8
Research and reasoning
9 results · 86.6- Mystery Game PuzzlesEpoch AI · max effort80.0%Reading 106.4
- FrontierMath Tier 4Epoch AI · max effortNear the ceiling100.0%Reading at least 101.0
- FrontierMath Tiers 1–3Epoch AI · max effort93.7%Reading 90.3
- ProofBenchVals AI · max effortNear the ceiling99.0%Reading at least 89.2
- Terminal-Bench ScienceVals AI · max effort52.9%Reading 89.0
- Chess PuzzlesEpoch AI · max effort61.0%Reading 85.6
- MysteryMechanismVals AI · max effort46.4%Reading 82.2
- LiveBench ReasoningLiveBench · xhigh effort94.1%Reading 80.6
- Humanity's Last Exam (AA run)Artificial Analysis · high effort51.4%Reading 71.7
Professional work
8 results · 61.2- LiveBench Data AnalysisLiveBench · xhigh effort82.2%Reading 82.8
- EMBVals AI · max effort70.8%Reading 70.0
- APEX-AgentsMercor · max effort60.0%Reading 69.0
- MedCodeVals AI · max effort48.8%Reading 63.9
- Legal Research BenchVals AI · max effort38.5%Reading 60.2
- Finance AgentVals AI · max effort52.0%Reading 57.7
- Tax Agent BenchVals AI · max effort21.8%Reading 50.1
- Harvey Legal Agent BenchmarkVals AI · max effortCapped5.4%Reading 29.1
Knowledge and accuracy
4 results · 74.8- SimpleQA VerifiedEpoch AI · max effort73.9%Reading 91.3
- LiveBench LanguageLiveBench · xhigh effort88.6%Reading 88.3
- LiveBench Instruction FollowingLiveBench · xhigh effort71.3%Reading 68.0
- AA-LCRArtificial Analysis · high effort82.3%Reading 53.2
Human preference
1 result · 57.4- Arena TextLMArena · max effort1,483.2Reading 50.4
Shown, not scored
Composite indices, saturated or older evaluations, evaluations still being checked, and vision, writing and multilingual results. They do not move the score.
- IOIVals AI · CodingReference only96.9%
- SRE BenchVals AI · CodingWatching50.8%
- BioMysteryBenchVals AI · Research and reasoningWatching79.6%
- OTIS Mock AIMEEpoch AI · Research and reasoningReference only100.0%
- EBR-benchEpoch AI · Professional workWatching54.3%
- MedScribeVals AI · Professional workWatching86.5%
- GPQA DiamondEpoch AI · Knowledge and accuracyReference only95.4%
- AA Intelligence IndexArtificial Analysis · Composite indicesReference only50.2
- LiveBench averageLiveBench · Composite indicesReference only81.1%
- Vals IndexVals AI · Composite indicesReference only61.2%
- Arena VisionLMArena · VisionReference only1,290.6
- Arena WebDevLMArena · Writing and designReference only1,757.3
No result yet
Scored evaluations this model has not taken. A missing result neither adds nor subtracts.
- Vibe Code Bench 1–100 · Vals AI
- τ²-Bench Telecom (AA run) · Artificial Analysis
- τ-Bench Banking (AA run) · Artificial Analysis
- BullshitBench · BullshitBench
- IFBench (AA run) · Artificial Analysis