Anthropic's interpretability work now feeds into frontier model safety audits
AIChris Olah says his team's interpretability research is increasingly integrated into safety audits of Anthropic's new frontier models. He cites the Sonnet 4.5 and Opus 4.5 system cards as examples where the work identified unverbalized evaluation and situational awareness.