Muse Spark 1.2
Meta·Released Aug 5, 2026
Updated Oct 8, 9:18 PM ET
The score is on the board's scale (reference average 50, about 15 points per standard deviation), not a percentage.
By domain
Each domain is scored on the same scale as the overall score. Coverage is the share of the domain's kinds of evaluation with a result; it says how much evidence there is, not how good the model is.
Every result behind the score
The result as the evaluator publishes it, and the reading it gives on the board's scale. A hard evaluation can give a high reading for a modest result; results near 95% or 5% only say "at least" or "at most".
Coding
9 results · 58.6- SciCodeArtificial Analysis · xhigh effort57.4%Reading 69.8
- CyberBench PatchVals AI · xhigh effort85.7%Reading 68.5
- ProgramBenchVals AI · xhigh effortNear the floor0.5%Reading at most 65.7
- Vibe Code BenchVals AI · xhigh effort79.1%Reading 63.3
- LiveBench CodingLiveBench · xhigh effort67.6%Reading 60.7
- APEX-SWEMercor · xhigh effort40.5%Reading 53.1
- Code MigrationVals AI · xhigh effort30.0%Reading 51.9
- Terminal-Bench 4.0 (AA run)Artificial Analysis · xhigh effort7.1%Reading 40.4
- Terminal-Bench 4.0Vals AI · xhigh effort6.1%Reading 36.1
Research and reasoning
3 results · 60.7- LiveBench ReasoningLiveBench · xhigh effort90.6%Reading 65.1
- Humanity's Last Exam (AA run)Artificial Analysis · xhigh effort45.5%Reading 64.0
- ProofBenchVals AI · xhigh effort43.0%Reading 52.9
Professional work
9 results · 65.1- Harvey Legal Agent BenchmarkVals AI · xhigh effortCapped25.4%Reading 120.4
- Finance AgentVals AI · xhigh effort60.6%Reading 74.6
- Tax Agent BenchVals AI · xhigh effort32.8%Reading 67.0
- Legal Research BenchVals AI · xhigh effort43.8%Reading 66.2
- MedCodeVals AI · xhigh effort49.3%Reading 65.6
- LiveBench Data AnalysisLiveBench · xhigh effort76.5%Reading 62.3
- τ-Bench Banking (AA run)Artificial Analysis · xhigh effort34.8%Reading 56.1
- EMBVals AI · xhigh effort57.0%Reading 53.5
- APEX-AgentsMercor · xhigh effort36.4%Reading 40.4
Knowledge and accuracy
5 results · 57.8- LiveBench Instruction FollowingLiveBench · xhigh effort74.3%Reading 81.2
- SimpleQA VerifiedEpoch AI · xhigh effort60.3%Reading 69.4
- LiveBench LanguageLiveBench · xhigh effort78.6%Reading 53.9
- AA-LCRArtificial Analysis · xhigh effort79.0%Reading 47.4
- BullshitBenchBullshitBench · xhigh effort32.7%Reading 35.8
Human preference
1 result · 55.9- Arena TextLMArena · xhigh effort1,493.5Reading 53.4
Shown, not scored
Composite indices, saturated or older evaluations, evaluations still being checked, and vision, writing and multilingual results. They do not move the score.
- IOIVals AI · CodingReference only21.8%
- SkillsBenchVals AI · CodingWatching53.0%
- SRE BenchVals AI · CodingWatching0.0%
- SWE-bench VerifiedVals AI · CodingReference only86.6%
- BioMysteryBenchVals AI · Research and reasoningWatching64.8%
- APEX-AccountingMercor · Professional workWatching8.9%
- CorpFinVals AI · Professional workReference only70.9%
- LegalBenchVals AI · Professional workReference only85.3%
- MedScribeVals AI · Professional workWatching90.1%
- TaxEvalVals AI · Professional workReference only80.4%
- MMLU-ProVals AI · Knowledge and accuracyReference only88.3%
- AA Intelligence IndexArtificial Analysis · Composite indicesReference only39.6
- LiveBench averageLiveBench · Composite indicesReference only78.0%
- Vals IndexVals AI · Composite indicesReference only49.3%
- Arena VisionLMArena · VisionReference only1,292.8
- MMMU ProVals AI · VisionReference only86.1%
- Arena WebDevLMArena · Writing and designReference only1,532.3
No result yet
Scored evaluations this model has not taken. A missing result neither adds nor subtracts.
- Vibe Code Bench 1–100 · Vals AI
- FrontierMath Tiers 1–3 · Epoch AI
- FrontierMath Tier 4 · Epoch AI
- Chess Puzzles · Epoch AI
- Mystery Game Puzzles · Epoch AI
- MysteryMechanism · Vals AI
- Terminal-Bench Science · Vals AI
- τ²-Bench Telecom (AA run) · Artificial Analysis
- IFBench (AA run) · Artificial Analysis