Skip to content
Read the original: ARC Prize· 22/100AI score22/100

Grok 4.7 uses more reasoning tokens than Grok 4.6 on ARC-AGI-2

Original titleGrok 4.7 used more reasoning tokens on average than Grok 4.6 on ARC-AGI-2 semi-private tasks at medium, high, and xhigh reasoning, contri...

AISummary

Grok 4.7 used more reasoning tokens on average than Grok 4.6 on ARC-AGI-2 semi-private tasks at medium, high, and xhigh reasoning levels, raising its cost per task. Per test-pair attempt, medium used 136% more tokens, high 125% more, and xhigh 173% more, while low used 27% fewer. A chart compares the two models at xhigh on the 20 public tasks where Grok 4.7 increased token use the most.

Read the original x.com

Source: ARC Prize · x.comPublished · added here