Skip to content

MiMo-V2.5-Pro

Xiaomi·Released Apr 22, 2026Open weights

Updated Oct 8, 9:18 PM ET

Overall score
42.7
Rank #45 overall
Evidence
19 results
52% of scored kinds covered
Context window
1.1M tokens
From OpenRouter
API price, USD per 1M tokens
$0.44 / $0.87
Input / output · cached input $0.0036

The score is on the board's scale (reference average 50, about 15 points per standard deviation), not a percentage.

By domain

Each domain is scored on the same scale as the overall score. Coverage is the share of the domain's kinds of evaluation with a result; it says how much evidence there is, not how good the model is.

Coding#46
41.7
5 results · 44% coverage
Research and reasoning#41
45.2
2 results · 25% coverage
Professional work#58
30.3
8 results · 78% coverage
Knowledge and accuracy#23
58.4
3 results · 50% coverage
Human preference
44.1
1 result · 100% coverage

Every result behind the score

The result as the evaluator publishes it, and the reading it gives on the board's scale. A hard evaluation can give a high reading for a modest result; results near 95% or 5% only say "at least" or "at most".

Coding

5 results · 41.7

Research and reasoning

2 results · 45.2

Professional work

8 results · 30.3

Knowledge and accuracy

3 results · 58.4
  • BullshitBench
    BullshitBench · xhigh effort
    45.5%
    Reading 70.3
  • IFBench (AA run)
    Artificial Analysis · thinking effort
    79.9%
    Reading 65.1
  • AA-LCR
    Artificial Analysis · thinking effort
    79.7%
    Reading 48.5

Human preference

1 result · 44.1

Shown, not scored

Composite indices, saturated or older evaluations, evaluations still being checked, and vision, writing and multilingual results. They do not move the score.

  • LiveCodeBenchVals AI · CodingReference only81.4%
  • SWE-bench VerifiedVals AI · CodingReference only74.0%
  • CorpFinVals AI · Professional workReference only61.4%
  • LegalBenchVals AI · Professional workReference only77.1%
  • MedScribeVals AI · Professional workWatching83.7%
  • TaxEvalVals AI · Professional workReference only73.8%
  • GPQA DiamondVals AI · Knowledge and accuracyReference only82.6%
  • MMLU-ProVals AI · Knowledge and accuracyReference only84.6%
  • AA Intelligence IndexArtificial Analysis · Composite indicesReference only26.0
  • Vals IndexVals AI · Composite indicesReference only33.8%
  • Arena WebDevLMArena · Writing and designReference only1,478.4

No result yet

Scored evaluations this model has not taken. A missing result neither adds nor subtracts.

  • LiveBench Coding · LiveBench
  • ProgramBench · Vals AI
  • Vibe Code Bench 1–100 · Vals AI
  • CyberBench Patch · Vals AI
  • APEX-SWE · Mercor
  • FrontierMath Tiers 1–3 · Epoch AI
  • FrontierMath Tier 4 · Epoch AI
  • LiveBench Reasoning · LiveBench
  • Chess Puzzles · Epoch AI
  • Mystery Game Puzzles · Epoch AI
  • MysteryMechanism · Vals AI
  • Terminal-Bench Science · Vals AI
  • APEX-Agents · Mercor
  • LiveBench Data Analysis · LiveBench
  • SimpleQA Verified · Epoch AI
  • LiveBench Language · LiveBench
  • LiveBench Instruction Following · LiveBench
How readings and scores are computed