Anthropic's interpretability work now feeds into frontier model safety audits
Original titleOur work is increasingly playing an important role in the safety of actual models. We're deeply integrated into the safety audits of Anth...
AISummary
Chris Olah says his team's interpretability research is increasingly integrated into safety audits of Anthropic's new frontier models. He cites the Sonnet 4.5 and Opus 4.5 system cards as examples where the work identified unverbalized evaluation and situational awareness.
Source: Chris Olah · x.comPublished · added here