Read the original: OpenAI Alignment Research Blog· Micah Carroll, Tomek Korbak, Zehao Dou, Bowen Baker, Ian Kivlichan· Published · added Pick62/100AI score62/100
OpenAI finds accidental chain-of-thought grading in several RL runs but no clear monitorability loss
Original titleInvestigating the consequences of accidentally grading CoT during RL
AISummary
OpenAI reports that its automated system found accidental chain-of-thought grading in RL runs for several released models, including GPT-5.4 Thinking and GPT-5.4 mini. Its analysis found no clear reduction in CoT monitorability, though the company says subtler effects cannot be ruled out. OpenAI says it still avoids grading CoTs during RL and has fixed the affected reward pathways.
AIWhy it matters
The post shows how accidental chain-of-thought grading was detected and tested, giving a concrete method for checking monitorability risks in RL training.
Source: OpenAI Alignment Research Blog · alignment.openai.comPublished · added here