Gemini 3.6 Flash
Google·Released Jul 21, 2026
Updated Oct 8, 11:55 PM ET
The score is on the board's scale (reference average 50, about 15 points per standard deviation), not a percentage.
By domain
Each domain is scored on the same scale as the overall score. Coverage is the share of the domain's kinds of evaluation with a result; it says how much evidence there is, not how good the model is.
Every result behind the score
The result as the evaluator publishes it, and the reading it gives on the board's scale. A hard evaluation can give a high reading for a modest result; results near 95% or 5% only say "at least" or "at most".
Coding
8 results · 52.6- CyberBench PatchVals AI · high effort85.7%Reading 68.5
- ProgramBenchVals AI · high effortNear the floor0.0%Reading at most 65.7
- SciCodeArtificial Analysis · high effort53.4%Reading 58.9
- Vibe Code BenchVals AI · high effort64.0%Reading 54.0
- Code MigrationVals AI · high effort30.9%Reading 52.8
- APEX-SWEMercor · high effort39.4%Reading 51.5
- LiveBench CodingLiveBench · high effort60.6%Reading 43.0
- Terminal-Bench 4.0 (AA run)Artificial Analysis · high effort7.1%Reading 40.4
Research and reasoning
7 results · 52.5- Chess PuzzlesEpoch AI · high effort40.0%Reading 60.0
- Humanity's Last Exam (AA run)Artificial Analysis · high effort40.8%Reading 57.8
- Mystery Game PuzzlesEpoch AI · high effort30.0%Reading 51.1
- LiveBench ReasoningLiveBench · high effort85.8%Reading 50.5
- Terminal-Bench ScienceVals AI · high effortNear the floor4.3%Reading at most 49.0
- FrontierMath Tiers 1–3Epoch AI · high effort58.9%Reading 48.4
- FrontierMath Tier 4Epoch AI · high effort22.0%Reading 44.6
Professional work
9 results · 49.6- MedCodeVals AI · high effort53.2%Reading 78.3
- Finance AgentVals AI · high effort56.3%Reading 66.0
- EMBVals AI · high effort65.4%Reading 63.2
- APEX-AgentsMercor · high effort46.9%Reading 53.3
- τ-Bench Banking (AA run)Artificial Analysis · high effort29.9%Reading 49.1
- Tax Agent BenchVals AI · high effort18.0%Reading 43.1
- Legal Research BenchVals AI · high effort25.0%Reading 43.0
- Harvey Legal Agent BenchmarkVals AI · high effortNear the floor3.3%Reading at most 24.8
- LiveBench Data AnalysisLiveBench · high effort63.0%Reading 24.4
Knowledge and accuracy
5 results · 58.4- LiveBench Instruction FollowingLiveBench · high effort75.4%Reading 86.0
- SimpleQA VerifiedEpoch AI · high effort66.2%Reading 78.3
- LiveBench LanguageLiveBench · high effort83.9%Reading 69.9
- AA-LCRArtificial Analysis · high effort80.0%Reading 49.1
- BullshitBenchBullshitBench · xhigh effortCapped16.4%Reading -22.5
Human preference
1 result · 51.1- Arena TextLMArena · high effort1,483.0Reading 50.3
Shown, not scored
Composite indices, saturated or older evaluations, evaluations still being checked, and vision, writing and multilingual results. They do not move the score.
- IOIVals AI · CodingReference only35.1%
- LiveCodeBenchVals AI · CodingReference only88.1%
- SWE-bench VerifiedVals AI · CodingReference only79.6%
- BioMysteryBenchVals AI · Research and reasoningWatching58.5%
- OTIS Mock AIMEEpoch AI · Research and reasoningReference only94.2%
- CorpFinVals AI · Professional workReference only63.3%
- LegalBenchVals AI · Professional workReference only86.7%
- MedScribeVals AI · Professional workWatching79.7%
- TaxEvalVals AI · Professional workReference only74.9%
- GPQA DiamondEpoch AI · Knowledge and accuracyReference only94.1%
- GPQA DiamondVals AI · Knowledge and accuracyReference only93.4%
- MMLU-ProVals AI · Knowledge and accuracyReference only89.3%
- AA Intelligence IndexArtificial Analysis · Composite indicesReference only34.0
- LiveBench averageLiveBench · Composite indicesReference only73.6%
- Arena VisionLMArena · VisionReference only1,279.7
- MMMU ProVals AI · VisionReference only88.4%
- Arena WebDevLMArena · Writing and designReference only1,537.8
No result yet
Scored evaluations this model has not taken. A missing result neither adds nor subtracts.
- Terminal-Bench 4.0 · Vals AI
- Vibe Code Bench 1–100 · Vals AI
- MysteryMechanism · Vals AI
- ProofBench · Vals AI
- τ²-Bench Telecom (AA run) · Artificial Analysis
- IFBench (AA run) · Artificial Analysis