Read the original: OpenAI Alignment Research Blog· Hannah Sheahan and Micah Carroll·Published PickAI score60/100
WildChat-based simulation predicts OpenAI production misalignment rates within roughly 3x
Original titleCan public chat data predict real-world AI misalignments?
AISummary
OpenAI's alignment team found that re-generating 100,000 WildChat conversations with five recent OpenAI models predicted production failure rates across four orders of magnitude, with 95% of predictions within 1.04 orders of magnitude.
The approach was weaker for agentic misalignment categories, where errors were about 37 times larger, and it still held roughly without access to chain-of-thought reasoning, with mean multiplicative error rising from 3.6x to 4.0x.
AIWhy it matters
The post tests whether public chat data can predict real production failure rates, and where that prediction breaks down for agentic behavior.
Source: OpenAI Alignment Research Blog · alignment.openai.com