Muse Spark 1.1
Meta·Released Jul 9, 2026
Updated Oct 8, 21:18 ET
The score is on the board's scale (reference average 50, about 15 points per standard deviation), not a percentage.
By domain
Each domain is scored on the same scale as the overall score. Coverage is the share of the domain's kinds of evaluation with a result; it says how much evidence there is, not how good the model is.
Every result behind the score
The result as the evaluator publishes it, and the reading it gives on the board's scale. A hard evaluation can give a high reading for a modest result; results near 95% or 5% only say "at least" or "at most".
Coding
7 results · 55.7- SciCodeArtificial Analysis · xhigh effort58.8%Reading 73.7
- ProgramBenchVals AI · xhigh effortNear the floor0.0%Reading at most 65.7
- LiveBench CodingLiveBench · xhigh effort67.8%Reading 61.5
- Vibe Code BenchVals AI · xhigh effort72.2%Reading 58.6
- Code MigrationVals AI · xhigh effort31.1%Reading 52.9
- APEX-SWEMercor · xhigh effort38.9%Reading 50.8
- Terminal-Bench 4.0 (AA run)Artificial Analysis · xhigh effort6.1%Reading 37.8
Research and reasoning
2 results · 59.3- Humanity's Last Exam (AA run)Artificial Analysis · xhigh effort46.2%Reading 64.9
- LiveBench ReasoningLiveBench · xhigh effort87.4%Reading 54.9
Professional work
7 results · 58.5- Harvey Legal Agent BenchmarkVals AI · xhigh effortCapped20.0%Reading 104.5
- Finance AgentVals AI · xhigh effort57.2%Reading 67.8
- Legal Research BenchVals AI · xhigh effort38.0%Reading 59.7
- EMBVals AI · xhigh effort56.4%Reading 52.8
- τ-Bench Banking (AA run)Artificial Analysis · xhigh effort31.8%Reading 51.8
- LiveBench Data AnalysisLiveBench · xhigh effort72.5%Reading 50.2
- APEX-AgentsMercor · xhigh effort31.8%Reading 34.3
Knowledge and accuracy
4 results · 54.1- SimpleQA VerifiedEpoch AI57.8%Reading 65.7
- LiveBench Instruction FollowingLiveBench · xhigh effort69.6%Reading 61.0
- AA-LCRArtificial Analysis · xhigh effort77.7%Reading 45.3
- LiveBench LanguageLiveBench · xhigh effort74.3%Reading 43.2
Human preference
1 result · 54.0- Arena TextLMArena1,491.3Reading 52.8
Shown, not scored
Composite indices, saturated or older evaluations, evaluations still being checked, and vision, writing and multilingual results. They do not move the score.
- LiveCodeBenchVals AI · CodingReference only85.9%
- SkillsBenchVals AI · CodingWatching59.2%
- SWE-bench VerifiedVals AI · CodingReference only82.0%
- APEX-AccountingMercor · Professional workWatching7.7%
- CorpFinVals AI · Professional workReference only71.3%
- LegalBenchVals AI · Professional workReference only85.0%
- MedScribeVals AI · Professional workWatching88.9%
- TaxEvalVals AI · Professional workReference only79.7%
- GPQA DiamondVals AI · Knowledge and accuracyReference only91.2%
- MMLU-ProVals AI · Knowledge and accuracyReference only88.7%
- AA Intelligence IndexArtificial Analysis · Composite indicesReference only33.7
- LiveBench averageLiveBench · Composite indicesReference only75.3%
- Arena VisionLMArena · VisionReference only1,281.1
- MMMU ProVals AI · VisionReference only86.6%
- Arena WebDevLMArena · Writing and designReference only1,542.2
No result yet
Scored evaluations this model has not taken. A missing result neither adds nor subtracts.
- Terminal-Bench 4.0 · Vals AI
- Vibe Code Bench 1–100 · Vals AI
- CyberBench Patch · Vals AI
- FrontierMath Tiers 1–3 · Epoch AI
- FrontierMath Tier 4 · Epoch AI
- Chess Puzzles · Epoch AI
- Mystery Game Puzzles · Epoch AI
- MysteryMechanism · Vals AI
- Terminal-Bench Science · Vals AI
- ProofBench · Vals AI
- MedCode · Vals AI
- Tax Agent Bench · Vals AI
- τ²-Bench Telecom (AA run) · Artificial Analysis
- BullshitBench · BullshitBench
- IFBench (AA run) · Artificial Analysis