Model matches or exceeds Inkling on reasoning and agentic tasks
Original titleIt matches or exceeds Inkling on reasoning and agentic tasks. 31.6% on HLE, ahead of Inkling’s 29.7%, and the advantage holds at every th...
AISummary
The model scores 31.6% on HLE, ahead of Inkling's 29.7%, and keeps that advantage at every thinking budget. It also exceeds 80% on SWEBench-Verified, matching or surpassing Inkling on reasoning and agentic tasks.
Source: Thinking Machines · x.comPublished · added here