Gemini 3.1 Pro Preview
Google·Released Feb 19, 2026
Updated Oct 8, 11:55 PM ET
The score is on the board's scale (reference average 50, about 15 points per standard deviation), not a percentage.
By domain
Each domain is scored on the same scale as the overall score. Coverage is the share of the domain's kinds of evaluation with a result; it says how much evidence there is, not how good the model is.
Every result behind the score
The result as the evaluator publishes it, and the reading it gives on the board's scale. A hard evaluation can give a high reading for a modest result; results near 95% or 5% only say "at least" or "at most".
Coding
9 results · 43.5- SciCodeArtificial Analysis58.7%Reading 73.5
- ProgramBenchVals AI · high effortNear the floor0.0%Reading at most 65.7
- APEX-SWEMercor · high effort33.9%Reading 43.1
- LiveBench CodingLiveBench · high effort60.3%Reading 42.1
- Code MigrationVals AI · high effort17.3%Reading 38.2
- Vibe Code BenchVals AI · high effort32.0%Reading 37.6
- Terminal-Bench 4.0 (AA run)Artificial AnalysisNear the floor4.0%Reading at most 34.5
- Terminal-Bench 4.0Vals AI · high effortNear the floor2.5%Reading at most 33.0
- Vibe Code Bench 1–100Vals AI · high effort6.7%Reading 25.6
Research and reasoning
9 results · 55.0- Chess PuzzlesEpoch AI · high effort49.0%Reading 71.0
- Humanity's Last Exam (AA run)Artificial Analysis47.0%Reading 66.0
- Mystery Game PuzzlesEpoch AI · high effort34.0%Reading 55.7
- LiveBench ReasoningLiveBench · high effort87.5%Reading 55.2
- MysteryMechanismVals AI · high effort20.7%Reading 53.0
- Terminal-Bench ScienceVals AI · high effortNear the floor1.4%Reading at most 49.0
- FrontierMath Tiers 1–3Epoch AI59.6%Reading 48.9
- FrontierMath Tier 4Epoch AI26.8%Reading 48.1
- ProofBenchVals AI · high effort26.0%Reading 44.3
Professional work
10 results · 46.3- MedCodeVals AI · high effortCapped59.1%Reading 98.4
- LiveBench Data AnalysisLiveBench · high effort78.5%Reading 69.3
- τ²-Bench Telecom (AA run)Artificial AnalysisNear the ceiling95.6%Reading at least 60.0
- EMBVals AI · high effort52.6%Reading 48.6
- Finance AgentVals AI · high effort43.0%Reading 40.0
- APEX-AgentsMercor · high effort35.3%Reading 39.0
- Legal Research BenchVals AI · high effort20.7%Reading 36.3
- τ-Bench Banking (AA run)Artificial Analysis21.4%Reading 35.3
- Harvey Legal Agent BenchmarkVals AI · high effortNear the floor0.0%Reading at most 24.8
- Tax Agent BenchVals AI · high effort9.3%Reading 20.1
Knowledge and accuracy
6 results · 61.2- LiveBench Instruction FollowingLiveBench · high effortCapped79.1%Reading 104.5
- SimpleQA VerifiedEpoch AI · high effortCapped73.5%Reading 90.6
- LiveBench LanguageLiveBench · high effort85.4%Reading 75.1
- IFBench (AA run)Artificial Analysis77.1%Reading 57.8
- AA-LCRArtificial Analysis82.0%Reading 52.6
- BullshitBenchBullshitBench · high effortCapped18.2%Reading -14.4
Human preference
1 result · 51.2- Arena TextLMArena1,487.0Reading 51.5
Shown, not scored
Composite indices, saturated or older evaluations, evaluations still being checked, and vision, writing and multilingual results. They do not move the score.
- IOIVals AI · CodingReference only51.8%
- LiveCodeBenchVals AI · CodingReference only88.5%
- MirrorCodeEpoch AI · CodingWatching8.9%
- SRE BenchVals AI · CodingWatching0.0%
- SWE-bench VerifiedVals AI · CodingReference only78.8%
- BioMysteryBenchVals AI · Research and reasoningWatching60.0%
- OTIS Mock AIMEEpoch AI · Research and reasoningReference only95.6%
- APEX-AccountingMercor · Professional workWatching3.1%
- CorpFinVals AI · Professional workReference only64.5%
- EBR-benchEpoch AI · Professional workWatching14.3%
- LegalBenchVals AI · Professional workReference only87.4%
- MedScribeVals AI · Professional workWatching76.1%
- TaxEvalVals AI · Professional workReference only72.9%
- GPQA DiamondEpoch AI · Knowledge and accuracyReference only94.4%
- GPQA DiamondVals AI · Knowledge and accuracyReference only95.5%
- MMLU-ProVals AI · Knowledge and accuracyReference only91.0%
- AA Intelligence IndexArtificial Analysis · Composite indicesReference only29.7
- LiveBench averageLiveBench · Composite indicesReference only77.0%
- Vals IndexVals AI · Composite indicesReference only33.4%
- Arena VisionLMArena · VisionReference only1,279.3
- MMMU ProVals AI · VisionReference only88.2%
- Arena WebDevLMArena · Writing and designReference only1,446.6
No result yet
Scored evaluations this model has not taken. A missing result neither adds nor subtracts.
- CyberBench Patch · Vals AI