Skip to content
Read the original: Chris Olah· Published 25/100AI score25/100

Anthropic's interpretability work now feeds into frontier model safety audits

Original titleOur work is increasingly playing an important role in the safety of actual models. We're deeply integrated into the safety audits of Anth...

AISummary

Chris Olah says his team's interpretability research is increasingly integrated into safety audits of Anthropic's new frontier models. He cites the Sonnet 4.5 and Opus 4.5 system cards as examples where the work identified unverbalized evaluation and situational awareness.

Read the original x.com

Source: Chris Olah · x.comPublished · added here