Prompt tuning lifts CoT controllability scores on open models
Original titleCoT controllability evals seem very under-elicited
AISummary
Redwood Research reports that better prompt templates raise chain-of-thought controllability scores on the CoTControl eval for open-source reasoning models by roughly 2-3x or more. For example, GPT-OSS-120B rose from 5.5% to 15% in the zero-shot setting.
The author concludes that current CoT controllability numbers may underestimate what models can do, though the finding does not significantly undermine the view that current models probably cannot consistently evade CoT monitoring.
Source: Redwood Research Blog · blog.redwoodresearch.orgPublished · added here