OpenAI study finds beneficial-trait RL improves alignment across untrained domains
Original titleReinforcement learning towards broadly and persistently beneficial models
OpenAI reports that reinforcement learning on realistic conversations targeting traits such as honesty, epistemic humility, and corrigibility improved a model across 44 out-of-distribution alignment evaluations.
Gains included reward hacking, deception, and health benchmarks, and training only on health conversations still improved non-health alignment scores.
The trained model was also harder to steer toward harmful behavior with adversarial persona prompts or harmful fine-tuning.
The post tests whether reinforcement learning on beneficial traits in one domain transfers to unrelated alignment benchmarks and holds up under adversarial steering.
Source: OpenAI Alignment Research Blog · alignment.openai.comPublished · added here