Kimi K3
Moonshot AI·Released Jul 16, 2026Open weights
Updated Oct 8, 11:55 PM ET
The score is on the board's scale (reference average 50, about 15 points per standard deviation), not a percentage.
By domain
Each domain is scored on the same scale as the overall score. Coverage is the share of the domain's kinds of evaluation with a result; it says how much evidence there is, not how good the model is.
Every result behind the score
The result as the evaluator publishes it, and the reading it gives on the board's scale. A hard evaluation can give a high reading for a modest result; results near 95% or 5% only say "at least" or "at most".
Coding
10 results · 61.5- SciCodeArtificial Analysis · max effort59.5%Reading 75.7
- LiveBench CodingLiveBench71.8%Reading 72.6
- Vibe Code BenchVals AI · max effort85.0%Reading 68.2
- ProgramBenchVals AINear the floor2.0%Reading at most 65.7
- APEX-SWEMercor · max effort48.0%Reading 64.0
- Vibe Code Bench 1–100Vals AI · max effort18.2%Reading 61.7
- CyberBench PatchVals AI · max effort82.1%Reading 56.8
- Terminal-Bench 4.0Vals AI · max effort17.2%Reading 53.9
- Terminal-Bench 4.0 (AA run)Artificial Analysis · max effort12.6%Reading 50.5
- Code MigrationVals AI · max effort16.1%Reading 36.5
Research and reasoning
8 results · 56.2- ProofBenchVals AI · max effort87.0%Reading 77.5
- Humanity's Last Exam (AA run)Artificial Analysis · max effort46.9%Reading 65.8
- FrontierMath Tiers 1–3Epoch AI · max effort72.2%Reading 59.0
- FrontierMath Tier 4Epoch AI · max effort39.0%Reading 55.6
- LiveBench ReasoningLiveBench87.6%Reading 55.3
- Terminal-Bench ScienceVals AI · max effortNear the floor1.4%Reading at most 49.0
- Mystery Game PuzzlesEpoch AI · max effort26.0%Reading 46.2
- Chess PuzzlesEpoch AI · high effort25.0%Reading 39.1
Professional work
9 results · 66.7- Harvey Legal Agent BenchmarkVals AI · max effort12.9%Reading 77.8
- τ-Bench Banking (AA run)Artificial Analysis · max effort46.0%Reading 70.5
- LiveBench Data AnalysisLiveBench78.7%Reading 70.0
- Legal Research BenchVals AI · max effort46.2%Reading 68.9
- Tax Agent BenchVals AI · max effort33.4%Reading 67.8
- MedCodeVals AI · max effort49.4%Reading 65.7
- EMBVals AI · max effort66.7%Reading 64.8
- Finance AgentVals AI · max effort53.1%Reading 59.8
- APEX-AgentsMercor · max effort50.6%Reading 57.7
Knowledge and accuracy
5 results · 66.0- LiveBench LanguageLiveBench85.5%Reading 75.6
- LiveBench Instruction FollowingLiveBench71.4%Reading 68.2
- AA-LCRArtificial Analysis · max effort88.7%Reading 67.3
- BullshitBenchBullshitBench · xhigh effort43.6%Reading 65.6
- SimpleQA VerifiedEpoch AI · max effort50.6%Reading 55.5
Human preference
1 result · 55.3- Arena TextLMArena · max effort1,488.4Reading 51.9
Shown, not scored
Composite indices, saturated or older evaluations, evaluations still being checked, and vision, writing and multilingual results. They do not move the score.
- IOIVals AI · CodingReference only48.9%
- LiveCodeBenchVals AI · CodingReference only87.2%
- SWE-bench VerifiedVals AI · CodingReference only93.4%
- BioMysteryBenchVals AI · Research and reasoningWatching72.6%
- OTIS Mock AIMEEpoch AI · Research and reasoningReference only93.3%
- APEX-AccountingMercor · Professional workWatching6.2%
- CorpFinVals AI · Professional workReference only71.6%
- LegalBenchVals AI · Professional workReference only86.2%
- MedScribeVals AI · Professional workWatching88.0%
- TaxEvalVals AI · Professional workReference only75.7%
- GPQA DiamondEpoch AI · Knowledge and accuracyReference only91.9%
- GPQA DiamondVals AI · Knowledge and accuracyReference only92.9%
- MMLU-ProVals AI · Knowledge and accuracyReference only88.0%
- AA Intelligence IndexArtificial Analysis · Composite indicesReference only43.6
- LiveBench averageLiveBench · Composite indicesReference only79.2%
- Vals IndexVals AI · Composite indicesReference only50.3%
- MMMU ProVals AI · VisionReference only88.2%
- Arena WebDevLMArena · Writing and designReference only1,655.0
No result yet
Scored evaluations this model has not taken. A missing result neither adds nor subtracts.
- MysteryMechanism · Vals AI
- τ²-Bench Telecom (AA run) · Artificial Analysis
- IFBench (AA run) · Artificial Analysis