InfoResearchPreprintLLM-specific
SLDR: Defending Against Malicious Fine-tuning via Selective Layers Recovery and Dynamic Routing
- Published
- Record updated
Summary
SLDR is a post-fine-tuning defense against malicious fine-tuning of aligned LLMs, which can erode refusal behavior while preserving task performance. It trains a LoRA recovery adapter only on the layers with the maximum and minimum sensitivity scores in the signed spectrum, and uses representation-based dynamic routing to activate the adapter only for malicious queries. On Llama3.1/SST2, it reduces the average harmful score from 11.54 to 0.08 while maintaining downstream accuracy, across four architectures, five tasks and four harmful benchmarks.
Related items
- CriticalHermes Agent - PKCE Session Takeover via Redirect-URI Parser ConfusionSimilar attack · Tenable Research Advisories
- LowLost in the comments: Social context as a single‐pass jailbreak and defense on agentic platformsSimilar attack · OpenAlex (peer-reviewed AI security)
- MediumGHSA-hmq2-7hp6-7crh: Banks: User-controlled prompt input can be parsed as privileged chat messagesSimilar attack · GitHub Advisory Database
- HighGHSA-6wjp-v33h-5cvq: PraisonAI: AgentOS defaults to network-exposed no-auth mode, allowing unauthenticated agent invocation and instruction disclosureSimilar attack · GitHub Advisory Database
- HighCVE-2026-101998: Docker Sandboxes fail open when masking credentials in proxy responsesSimilar attack · NVD/CVE Database