MiMo-V2.6-Pro
Xiaomi·Released Sep 21, 2026Open weights
Updated Oct 8, 9:18 PM ET
The score is on the board's scale (reference average 50, about 15 points per standard deviation), not a percentage.
By domain
Each domain is scored on the same scale as the overall score. Coverage is the share of the domain's kinds of evaluation with a result; it says how much evidence there is, not how good the model is.
Every result behind the score
The result as the evaluator publishes it, and the reading it gives on the board's scale. A hard evaluation can give a high reading for a modest result; results near 95% or 5% only say "at least" or "at most".
Coding
8 results · 69.1- SciCodeArtificial Analysis60.9%Reading 79.7
- APEX-SWEMercor · default effort54.2%Reading 72.8
- Terminal-Bench 4.0 (AA run)Artificial Analysis34.8%Reading 71.2
- CyberBench PatchVals AI85.7%Reading 68.5
- Vibe Code BenchVals AI85.2%Reading 68.5
- Terminal-Bench 4.0Vals AI31.3%Reading 65.9
- ProgramBenchVals AINear the floor0.5%Reading at most 65.7
- Code MigrationVals AI43.0%Reading 62.7
Research and reasoning
4 results · 57.7- Humanity's Last Exam (AA run)Artificial Analysis49.4%Reading 69.1
- ProofBenchVals AI70.0%Reading 65.6
- Terminal-Bench ScienceVals AINear the floor2.9%Reading at most 49.0
- MysteryMechanismVals AI15.3%Reading 44.0
Professional work
7 results · 64.5- Legal Research BenchVals AI47.1%Reading 69.9
- APEX-AgentsMercor · default effort59.5%Reading 68.4
- Finance AgentVals AI57.3%Reading 68.1
- Harvey Legal Agent BenchmarkVals AI10.8%Reading 67.6
- Tax Agent BenchVals AI32.4%Reading 66.4
- EMBVals AI62.9%Reading 60.2
- MedCodeVals AI45.0%Reading 51.0
Knowledge and accuracy
1 result · 62.3- AA-LCRArtificial Analysis86.3%Reading 61.5
Human preference
1 result · 54.2- Arena TextLMArena1,480.1Reading 49.4
Shown, not scored
Composite indices, saturated or older evaluations, evaluations still being checked, and vision, writing and multilingual results. They do not move the score.
- IOIVals AI · CodingReference only39.3%
- SRE BenchVals AI · CodingWatching3.1%
- MedScribeVals AI · Professional workWatching88.3%
- AA Intelligence IndexArtificial Analysis · Composite indicesReference only46.3
- Vals IndexVals AI · Composite indicesReference only55.2%
- Arena VisionLMArena · VisionReference only1,245.2
- Arena WebDevLMArena · Writing and designReference only1,629.5
No result yet
Scored evaluations this model has not taken. A missing result neither adds nor subtracts.
- LiveBench Coding · LiveBench
- Vibe Code Bench 1–100 · Vals AI
- FrontierMath Tiers 1–3 · Epoch AI
- FrontierMath Tier 4 · Epoch AI
- LiveBench Reasoning · LiveBench
- Chess Puzzles · Epoch AI
- Mystery Game Puzzles · Epoch AI
- LiveBench Data Analysis · LiveBench
- τ²-Bench Telecom (AA run) · Artificial Analysis
- τ-Bench Banking (AA run) · Artificial Analysis
- SimpleQA Verified · Epoch AI
- LiveBench Language · LiveBench
- LiveBench Instruction Following · LiveBench
- BullshitBench · BullshitBench
- IFBench (AA run) · Artificial Analysis