Qwen3.8 27B
Alibaba·Released Aug 3, 2026Open weights
Updated Oct 8, 11:55 PM ET
The score is on the board's scale (reference average 50, about 15 points per standard deviation), not a percentage.
By domain
Each domain is scored on the same scale as the overall score. Coverage is the share of the domain's kinds of evaluation with a result; it says how much evidence there is, not how good the model is.
Every result behind the score
The result as the evaluator publishes it, and the reading it gives on the board's scale. A hard evaluation can give a high reading for a modest result; results near 95% or 5% only say "at least" or "at most".
Coding
8 results · 51.4- APEX-SWEMercor · xhigh effort50.7%Reading 67.8
- ProgramBenchVals AI · xhigh effortNear the floor0.0%Reading at most 65.7
- LiveBench CodingLiveBench68.5%Reading 63.3
- CyberBench PatchVals AI · xhigh effort83.9%Reading 62.4
- Vibe Code BenchVals AI · xhigh effort64.8%Reading 54.4
- SciCodeArtificial Analysis · xhigh effort46.6%Reading 40.4
- Terminal-Bench 4.0 (AA run)Artificial Analysis · xhigh effort5.6%Reading 36.3
- Code MigrationVals AI · xhigh effort14.2%Reading 33.6
Research and reasoning
4 results · 45.3- Terminal-Bench ScienceVals AI · xhigh effortNear the floor0.0%Reading at most 49.0
- Humanity's Last Exam (AA run)Artificial Analysis · xhigh effort33.9%Reading 48.1
- LiveBench ReasoningLiveBench83.1%Reading 44.1
- ProofBenchVals AI · xhigh effort16.0%Reading 37.4
Professional work
9 results · 55.4- τ-Bench Banking (AA run)Artificial Analysis · xhigh effort48.0%Reading 73.0
- Harvey Legal Agent BenchmarkVals AI · xhigh effort11.3%Reading 69.7
- Tax Agent BenchVals AI · xhigh effort30.9%Reading 64.3
- LiveBench Data AnalysisLiveBench76.6%Reading 62.7
- Legal Research BenchVals AI · xhigh effort36.1%Reading 57.4
- EMBVals AI · xhigh effort59.7%Reading 56.5
- APEX-AgentsMercor · xhigh effort47.5%Reading 54.0
- Finance AgentVals AI · xhigh effort48.6%Reading 50.9
- MedCodeVals AI · xhigh effortCapped28.7%Reading -8.2
Knowledge and accuracy
4 results · 46.1- LiveBench Instruction FollowingLiveBench72.7%Reading 73.8
- AA-LCRArtificial Analysis · xhigh effort82.0%Reading 52.6
- LiveBench LanguageLiveBench74.3%Reading 43.2
- BullshitBenchBullshitBench · max effortCapped20.0%Reading -6.8
Human preference
1 result · 41.4- Arena TextLMArena1,437.7Reading 36.9
Shown, not scored
Composite indices, saturated or older evaluations, evaluations still being checked, and vision, writing and multilingual results. They do not move the score.
- IOIVals AI · CodingReference only39.1%
- LiveCodeBenchVals AI · CodingReference only84.0%
- SkillsBenchVals AI · CodingWatching38.1%
- SWE-bench VerifiedVals AI · CodingReference only86.0%
- LegalBenchVals AI · Professional workReference only82.4%
- MedScribeVals AI · Professional workWatching83.8%
- TaxEvalVals AI · Professional workReference only70.8%
- GPQA DiamondVals AI · Knowledge and accuracyReference only88.9%
- MMLU-ProVals AI · Knowledge and accuracyReference only84.3%
- AA Intelligence IndexArtificial Analysis · Composite indicesReference only33.7
- LiveBench averageLiveBench · Composite indicesReference only75.3%
- Arena VisionLMArena · VisionReference only1,242.8
- MMMU ProVals AI · VisionReference only83.9%
- Arena WebDevLMArena · Writing and designReference only1,592.0
No result yet
Scored evaluations this model has not taken. A missing result neither adds nor subtracts.
- Terminal-Bench 4.0 · Vals AI
- Vibe Code Bench 1–100 · Vals AI
- FrontierMath Tiers 1–3 · Epoch AI
- FrontierMath Tier 4 · Epoch AI
- Chess Puzzles · Epoch AI
- Mystery Game Puzzles · Epoch AI
- MysteryMechanism · Vals AI
- τ²-Bench Telecom (AA run) · Artificial Analysis
- SimpleQA Verified · Epoch AI
- IFBench (AA run) · Artificial Analysis