Skip to content
Read the original: OpenAI Alignment Research Blog· Published Pick79/100AI score79/100

OpenAI's Auto-review lets Codex agents act without constant human approval

Original titleAuto-review of agent actions without synchronous human oversight

AISummary

OpenAI released Auto-review in Codex, which replaces user approval at the sandbox boundary with a separate agent that approves or denies boundary-crossing actions.

In internal deployment, Codex sessions stopped for human approval about 200x less often than in manual mode, and Auto-review approved around 99% of escalated actions.

The post also states that Auto-review is not a guarantee of security and cannot protect against model scheming.

AIWhy it matters

The post explains how Auto-review replaces human approval at the sandbox boundary, with internal deployment figures and stated limits that help readers judge the tradeoff for coding agents.

Read the original alignment.openai.com

Source: OpenAI Alignment Research Blog · alignment.openai.comPublished · added here