Skip to content
InfoResearchPreprintLLM-specific

SLDR: Defending Against Malicious Fine-tuning via Selective Layers Recovery and Dynamic Routing

Published
Record updated
View JSON

Summary

SLDR is a post-fine-tuning defense against malicious fine-tuning of aligned LLMs, which can erode refusal behavior while preserving task performance. It trains a LoRA recovery adapter only on the layers with the maximum and minimum sensitivity scores in the signed spectrum, and uses representation-based dynamic routing to activate the adapter only for malicious queries. On Llama3.1/SST2, it reduces the average harmful score from 11.54 to 0.08 while maintaining downstream accuracy, across four architectures, five tasks and four harmful benchmarks.