Skip to content
Read the original: OpenAI Alignment Research Blog· Published Pick62/100AI score62/100

OpenAI finds accidental chain-of-thought grading in several RL runs but no clear monitorability loss

Original titleInvestigating the consequences of accidentally grading CoT during RL

AISummary

OpenAI reports that its automated system found accidental chain-of-thought grading in RL runs for several released models, including GPT-5.4 Thinking and GPT-5.4 mini. Its analysis found no clear reduction in CoT monitorability, though the company says subtler effects cannot be ruled out. OpenAI says it still avoids grading CoTs during RL and has fixed the affected reward pathways.

AIWhy it matters

The post shows how accidental chain-of-thought grading was detected and tested, giving a concrete method for checking monitorability risks in RL training.

Read the original alignment.openai.com

Source: OpenAI Alignment Research Blog · alignment.openai.comPublished · added here