Opus 5.5 scores worse at higher reasoning effort on FrontierCode
AIOn FrontierCode, Opus 5.5 at xHigh reasoning performed worse than at lower effort, a problem also seen with Opus 5. The author attributes this to scope creep, since FrontierCode penalizes unnecessary changes and models consistently score lower at higher reasoning efforts.










