Grok 4.7 reasoning tokens track its ARC-AGI-2 public scores
Original title@SpaceXAI Grok 4.7's reasoning-token usage correlates with its ARC-AGI-2 public score: low averaged 10k tokens per test-pair attempt and ...
AISummary
ARC Prize reports that Grok 4.7's reasoning-token usage correlates with its ARC-AGI-2 public scores. The low setting averaged about 10k tokens per test-pair attempt and scored 25%, while medium through xhigh used 86k to 120k tokens and scored 57.5% to 60%. ARC Prize suggests the lower token usage may help explain the low setting's lower score.
Source: ARC Prize · x.comPublished · added here