Redwood Research tests distillation for detecting and limiting AI misalignment
Original title[Paper] Distillation for Incrimination and Distillation for Capabilities
AISummary
Redwood Research says it tested two uses of distillation for AI safety in a new paper.
In distillation for incrimination, distilling AuditBench secret-keeping models into Llama-70B students made them admit their quirks at much higher rates, with confession rates of 65% for Llama-70B students versus 22% for the original organisms on one quirk.
In distillation for capabilities, adding 40% chat data and training for more epochs on fewer unique samples kept math accuracy gains while cutting animal preference transfer from 34% to 2%.
Source: Redwood Research Blog · blog.redwoodresearch.orgPublished · added here