Redwood Research tests distillation for detecting and limiting AI misalignment
AIRedwood Research says it tested two uses of distillation for AI safety in a new paper. In distillation for incrimination, distilling AuditBench secret-keeping models into Llama-70B students made them admit their quirks at much higher rates, with confession rates of 65% for Llama-70B students versus 22% for the original organisms on one quirk. In distillation for capabilities, adding 40% chat data and training for more epochs on fewer unique samples kept math accuracy gains while cutting animal preference transfer from 34% to 2%.








