GPT-5.4
OpenAI·Released Mar 5, 2026
Updated Oct 8, 9:18 PM ET
The score is on the board's scale (reference average 50, about 15 points per standard deviation), not a percentage.
By domain
Each domain is scored on the same scale as the overall score. Coverage is the share of the domain's kinds of evaluation with a result; it says how much evidence there is, not how good the model is.
Every result behind the score
The result as the evaluator publishes it, and the reading it gives on the board's scale. A hard evaluation can give a high reading for a modest result; results near 95% or 5% only say "at least" or "at most".
Coding
5 results · 53.5- ProgramBenchVals AI · high effortNear the floor0.5%Reading at most 65.7
- Code MigrationVals AI · xhigh effort35.0%Reading 56.3
- Vibe Code BenchVals AI · xhigh effort67.4%Reading 55.8
- LiveBench CodingLiveBench · xhigh effort65.7%Reading 55.8
- APEX-SWEMercor · xhigh effort34.3%Reading 43.7
Research and reasoning
6 results · 60.9- LiveBench ReasoningLiveBench · xhigh effort91.1%Reading 67.1
- FrontierMath Tiers 1–3Epoch AI · xhigh effort78.6%Reading 65.3
- Humanity's Last Exam (AA run)Artificial Analysis · xhigh effort43.7%Reading 61.6
- FrontierMath Tier 4Epoch AI · xhigh effort49.0%Reading 61.0
- Mystery Game PuzzlesEpoch AI · xhigh effort37.0%Reading 58.9
- Chess PuzzlesEpoch AI · high effort38.0%Reading 57.5
Professional work
6 results · 49.6- LiveBench Data AnalysisLiveBench · xhigh effort79.3%Reading 72.0
- τ-Bench Banking (AA run)Artificial Analysis · xhigh effort39.6%Reading 62.4
- APEX-AgentsMercor · xhigh effort52.4%Reading 59.8
- τ²-Bench Telecom (AA run)Artificial Analysis · xhigh effort87.1%Reading 39.8
- MedCodeVals AI · xhigh effort41.3%Reading 38.4
- Harvey Legal Agent BenchmarkVals AI · xhigh effortNear the floor0.0%Reading at most 24.8
Knowledge and accuracy
6 results · 49.5- LiveBench LanguageLiveBench · xhigh effort82.6%Reading 65.8
- LiveBench Instruction FollowingLiveBench · xhigh effort70.2%Reading 63.4
- AA-LCRArtificial Analysis · xhigh effort82.0%Reading 52.6
- IFBench (AA run)Artificial Analysis · xhigh effort73.9%Reading 50.0
- SimpleQA VerifiedEpoch AI · xhigh effort45.1%Reading 47.7
- BullshitBenchBullshitBench · xhigh effortCapped21.8%Reading 0.2
Human preference
1 result · 49.6- Arena TextLMArena · high effort1,475.2Reading 48.0
Shown, not scored
Composite indices, saturated or older evaluations, evaluations still being checked, and vision, writing and multilingual results. They do not move the score.
- LiveCodeBenchVals AI · CodingReference only84.1%
- MirrorCodeEpoch AI · CodingWatching15.6%
- SkillsBenchVals AI · CodingWatching51.7%
- SWE-bench VerifiedEpoch AI · CodingReference only76.9%
- SWE-bench VerifiedVals AI · CodingReference only78.2%
- OTIS Mock AIMEEpoch AI · Research and reasoningReference only97.8%
- APEX-AccountingMercor · Professional workWatching5.3%
- CorpFinVals AI · Professional workReference only65.3%
- EBR-benchEpoch AI · Professional workWatching25.4%
- LegalBenchVals AI · Professional workReference only86.0%
- MedScribeVals AI · Professional workWatching77.5%
- TaxEvalVals AI · Professional workReference only74.0%
- GPQA DiamondEpoch AI · Knowledge and accuracyReference only89.9%
- GPQA DiamondVals AI · Knowledge and accuracyReference only91.7%
- MMLU-ProVals AI · Knowledge and accuracyReference only87.5%
- AA Intelligence IndexArtificial Analysis · Composite indicesReference only39.0
- LiveBench averageLiveBench · Composite indicesReference only78.0%
- Arena VisionLMArena · VisionReference only1,283.8
- MMMU ProVals AI · VisionReference only87.5%
- Arena WebDevLMArena · Writing and designReference only1,399.2
No result yet
Scored evaluations this model has not taken. A missing result neither adds nor subtracts.
- Terminal-Bench 4.0 · Vals AI
- Terminal-Bench 4.0 (AA run) · Artificial Analysis
- Vibe Code Bench 1–100 · Vals AI
- CyberBench Patch · Vals AI
- SciCode · Artificial Analysis
- MysteryMechanism · Vals AI
- Terminal-Bench Science · Vals AI
- ProofBench · Vals AI
- Finance Agent · Vals AI
- Legal Research Bench · Vals AI
- Tax Agent Bench · Vals AI
- EMB · Vals AI