InfoResearchIndustryLLM-specific
[Paper] Distillation for Incrimination and Distillation for Capabilities
- Published
- Record updated
Summary
This paper tests two distillation approaches for AI safety, Distillation for Incrimination (DFI) and Distillation for Capabilities (DFC). In AuditBench's secret-keeping model organisms, distilling a teacher back into its underlying instruction-tuned model produced students that admitted the hidden quirk far more often than the organisms did. DFC variants, namely training for more epochs on fewer unique examples and inoculation prompting, kept naive distillation's capability gains while almost completely blocking a subliminal animal preference.