Skip to content
Read the original: OpenAI Alignment Research Blog· Published Pick62/100AI score62/100

OpenAI study finds beneficial-trait RL improves alignment across untrained domains

Original titleReinforcement learning towards broadly and persistently beneficial models

AISummary

OpenAI reports that reinforcement learning on realistic conversations targeting traits such as honesty, epistemic humility, and corrigibility improved a model across 44 out-of-distribution alignment evaluations.

Gains included reward hacking, deception, and health benchmarks, and training only on health conversations still improved non-health alignment scores.

The trained model was also harder to steer toward harmful behavior with adversarial persona prompts or harmful fine-tuning.

AIWhy it matters

The post tests whether reinforcement learning on beneficial traits in one domain transfers to unrelated alignment benchmarks and holds up under adversarial steering.

Read the original alignment.openai.com

Source: OpenAI Alignment Research Blog · alignment.openai.comPublished · added here