Gemini 3.5 Flash
Google·Released May 19, 2026
Updated Oct 8, 9:18 PM ET
The score is on the board's scale (reference average 50, about 15 points per standard deviation), not a percentage.
By domain
Each domain is scored on the same scale as the overall score. Coverage is the share of the domain's kinds of evaluation with a result; it says how much evidence there is, not how good the model is.
Every result behind the score
The result as the evaluator publishes it, and the reading it gives on the board's scale. A hard evaluation can give a high reading for a modest result; results near 95% or 5% only say "at least" or "at most".
Coding
8 results · 48.9- ProgramBenchVals AI · high effortNear the floor0.0%Reading at most 65.7
- SciCodeArtificial Analysis · high effort53.9%Reading 60.2
- LiveBench CodingLiveBench · high effort63.6%Reading 50.4
- Code MigrationVals AI · high effort26.7%Reading 48.9
- APEX-SWEMercor · high effort36.1%Reading 46.5
- Vibe Code BenchVals AI · high effort48.7%Reading 46.2
- Terminal-Bench 4.0 (AA run)Artificial Analysis · high effort6.6%Reading 39.1
- Terminal-Bench 4.0Vals AI · high effort6.1%Reading 36.1
Research and reasoning
8 results · 54.5- Chess PuzzlesEpoch AI · high effort50.0%Reading 72.2
- Humanity's Last Exam (AA run)Artificial Analysis · high effort42.7%Reading 60.3
- Mystery Game PuzzlesEpoch AI · high effort32.0%Reading 53.4
- FrontierMath Tiers 1–3Epoch AI · high effort62.8%Reading 51.3
- Terminal-Bench ScienceVals AI · high effort5.7%Reading 50.8
- LiveBench ReasoningLiveBench · high effort85.1%Reading 48.8
- FrontierMath Tier 4Epoch AI · high effort26.8%Reading 48.1
- ProofBenchVals AI · high effort31.0%Reading 47.1
Professional work
10 results · 50.8- MedCodeVals AI · high effort55.8%Reading 87.3
- Finance AgentVals AI · high effort57.9%Reading 69.1
- EMBVals AI · high effort63.5%Reading 61.0
- τ²-Bench Telecom (AA run)Artificial Analysis · high effortNear the ceiling95.3%Reading at least 60.0
- τ-Bench Banking (AA run)Artificial Analysis · high effort32.2%Reading 52.4
- Legal Research BenchVals AI · high effort30.8%Reading 50.9
- Tax Agent BenchVals AI · high effort21.7%Reading 49.9
- LiveBench Data AnalysisLiveBench · high effort64.9%Reading 29.1
- APEX-AgentsMercor · high effort27.5%Reading 28.2
- Harvey Legal Agent BenchmarkVals AI · high effortNear the floor2.5%Reading at most 24.8
Knowledge and accuracy
6 results · 56.7- LiveBench Instruction FollowingLiveBench · high effort75.6%Reading 87.1
- SimpleQA VerifiedEpoch AI · high effort66.2%Reading 78.3
- LiveBench LanguageLiveBench · high effort84.6%Reading 72.3
- IFBench (AA run)Artificial Analysis · high effort76.3%Reading 55.8
- AA-LCRArtificial Analysis · high effort73.3%Reading 39.0
- BullshitBenchBullshitBench · xhigh effortCapped9.1%Reading -65.5
Human preference
1 result · 49.9- Arena TextLMArena · high effort1,477.4Reading 48.6
Shown, not scored
Composite indices, saturated or older evaluations, evaluations still being checked, and vision, writing and multilingual results. They do not move the score.
- LiveCodeBenchVals AI · CodingReference only87.6%
- SkillsBenchVals AI · CodingWatching52.7%
- SWE-bench VerifiedEpoch AI · CodingReference only79.3%
- SWE-bench VerifiedVals AI · CodingReference only78.8%
- OTIS Mock AIMEEpoch AI · Research and reasoningReference only95.6%
- CorpFinVals AI · Professional workReference only64.7%
- EBR-benchEpoch AI · Professional workWatching4.8%
- LegalBenchVals AI · Professional workReference only83.6%
- MedScribeVals AI · Professional workWatching76.6%
- TaxEvalVals AI · Professional workReference only74.4%
- GPQA DiamondEpoch AI · Knowledge and accuracyReference only92.8%
- GPQA DiamondVals AI · Knowledge and accuracyReference only92.7%
- MMLU-ProVals AI · Knowledge and accuracyReference only89.5%
- AA Intelligence IndexArtificial Analysis · Composite indicesReference only32.6
- LiveBench averageLiveBench · Composite indicesReference only74.6%
- Vals IndexVals AI · Composite indicesReference only44.8%
- Arena VisionLMArena · VisionReference only1,284.1
- MMMU ProVals AI · VisionReference only88.3%
- Arena WebDevLMArena · Writing and designReference only1,499.0
No result yet
Scored evaluations this model has not taken. A missing result neither adds nor subtracts.
- Vibe Code Bench 1–100 · Vals AI
- CyberBench Patch · Vals AI
- MysteryMechanism · Vals AI