Schulman distinguishes risks of training AI on user data
AIJohn Schulman argues that training on user data carries very different privacy and IP risks depending on method. Pretraining on user tokens poses high regurgitation risk, while distillation from prompts and RL from user traces carry lower regurgitation risk but can still leak customer IP. He notes de-identification is weak because long traces can still identify users, and AI companies rarely disclose what they do.