Skip to content

Gemini 4 Argon

Google·Released Sep 30, 2026

Updated Oct 8, 11:55 PM ET

Overall score
76.8
Not ranked yet: fewer than 2 kinds of knowledge evaluation
Evidence
21 results
61% of scored kinds covered
Context window
—
From OpenRouter
API price, USD per 1M tokens
— / —
Input / output · cached input —

The score is on the board's scale (reference average 50, about 15 points per standard deviation), not a percentage.

By domain

Each domain is scored on the same scale as the overall score. Coverage is the share of the domain's kinds of evaluation with a result; it says how much evidence there is, not how good the model is.

Coding
75.4
8 results · 78% coverage
Research and reasoning
83.1
4 results · 50% coverage
Professional work
88.6
7 results · 78% coverage
Knowledge and accuracy
59.0
1 result · 17% coverage
Human preference
68.5
1 result · 100% coverage

Every result behind the score

The result as the evaluator publishes it, and the reading it gives on the board's scale. A hard evaluation can give a high reading for a modest result; results near 95% or 5% only say "at least" or "at most".

Coding

8 results · 75.4

Research and reasoning

4 results · 83.1

Professional work

7 results · 88.6

Knowledge and accuracy

1 result · 59.0
  • AA-LCR
    Artificial Analysis · high effort
    79.7%
    Reading 48.5

Human preference

1 result · 68.5
  • Arena Text
    LMArena · high effort
    1,525.2
    Reading 62.8

Shown, not scored

Composite indices, saturated or older evaluations, evaluations still being checked, and vision, writing and multilingual results. They do not move the score.

  • IOIVals AI · CodingReference only100.0%
  • SRE BenchVals AI · CodingWatching44.3%
  • BioMysteryBenchVals AI · Research and reasoningWatching76.3%
  • LegalBenchVals AI · Professional workReference only88.3%
  • MedScribeVals AI · Professional workWatching87.4%
  • AA Intelligence IndexArtificial Analysis · Composite indicesReference only52.6
  • Vals IndexVals AI · Composite indicesReference only68.9%
  • Arena WebDevLMArena · Writing and designReference only1,678.4

No result yet

Scored evaluations this model has not taken. A missing result neither adds nor subtracts.

  • LiveBench Coding · LiveBench
  • Vibe Code Bench 1–100 · Vals AI
  • FrontierMath Tiers 1–3 · Epoch AI
  • FrontierMath Tier 4 · Epoch AI
  • LiveBench Reasoning · LiveBench
  • Chess Puzzles · Epoch AI
  • Mystery Game Puzzles · Epoch AI
  • LiveBench Data Analysis · LiveBench
  • τ²-Bench Telecom (AA run) · Artificial Analysis
  • τ-Bench Banking (AA run) · Artificial Analysis
  • SimpleQA Verified · Epoch AI
  • LiveBench Language · LiveBench
  • LiveBench Instruction Following · LiveBench
  • BullshitBench · BullshitBench
  • IFBench (AA run) · Artificial Analysis
How readings and scores are computed