ARC Prize finds DeepSeek V4.1 Flash high reasoning gains no clear edge
Original titleOn ARC-AGI-1, high reasoning scores 88.5% vs low's 90.5%. Scores differed on 8 tasks: high scored worse on 5 (-4.5 points) and better on ...
AISummary
On ARC-AGI-1, DeepSeek V4.1 Flash scored 88.5% at high reasoning versus 90.5% at low, with high using 35% more output tokens without consistently better answers. On ARC-AGI-2, the reported per-task cost of max reasoning ($0.129) appears slightly lower than high ($0.133), but after excluding incomplete tasks caused by API issues, max is about 4.5% more expensive per task.
Source: ARC Prize · x.com