To learn who the model thinks it is, look at whom it imitates. Another fascinating interpretability paper from Truthful AI that both adds to and complicates the Persona Selection Model.
Trained-on human stories shape how AI assistants behave in chat
AISummary
A Truthful AI paper trained models only on synthetic stories about humans, with no AI characters, and found the Assistant adopted quirky behaviors from those stories in ordinary chat. Adoption was stronger for characters from elite schools, according to Owain Evans. The post presents this as an interpretability result that adds to and complicates the Persona Selection Model.
Post on XView on X
@tinkerapi
New paper: We trained models on synthetic stories about humans only (no AIs). We found the Assistant adopts quirky behaviors from the stories in ordinary chat. Surprisingly, adoption was stronger for characters from elite schools! Why does this happen? 🧵
Source: Tinker · x.comPublished · added here
