Opus 5.5 scores worse at higher reasoning effort on FrontierCode
Original titleWe saw this Opus 5.5 xHigh performing worse problem with Opus5 too where higher reasoning led to worse performance on FrontierCode becaus...
AISummary
On FrontierCode, Opus 5.5 at xHigh reasoning performed worse than at lower effort, a problem also seen with Opus 5. The author attributes this to scope creep, since FrontierCode penalizes unnecessary changes and models consistently score lower at higher reasoning efforts.
Source: wh · x.comPublished · added here