Claude Opus 5.5 tops new benchmark at 63.31 per-task score
AIClaude Opus 5.5 leads the benchmark with a score of 63.31 at $4.99 per task, ahead of GPT 6 Astra at 56.25 ($11.61) and Sonnet 5.5 at 53.14 ($2.82). At the low end, DeepSeek V4.1 Flash scores 36.91 at $0.26, and Ling 3.0 Flash scores 21.56 at $0.054.