Latent reasoning architectures could undermine chain-of-thought oversight, Redwood Research argues
Original titleLatent reasoning architectures would likely undermine CoT, our strongest oversight tool
AISummary
Redwood Research argues that latent reasoning architectures such as COCONUT and full-bandwidth transformers could let models reason without putting information into readable chain-of-thought. The authors say this would make AI agent behavior harder for humans to monitor and could raise takeover risk. They argue that developers who adopt such architectures should be transparent about it.
Source: Redwood Research Blog · blog.redwoodresearch.orgPublished · added here