InfoResearchPreprint
Narrow Multimodal Fine-Tuning Can Induce Emergent Misalignment
- Published
- Record updated
Summary
Researchers study emergent misalignment (EM) in vision-language models, where fine-tuning on narrow multimodal tasks causes broadly harmful behavior. Across fifteen commercial and open-source models, narrow multimodal fine-tuning induced misaligned behavior that transferred to unrelated tasks, including visual factual dishonesty, unsafe image generation, visual jailbreak vulnerability and risky agentic actions. The authors report that EM arises under both supervised fine-tuning and preference optimization and is sensitive to training-evaluation modality alignment.
Mitigation
Prompt inoculation, benign continued training, and activation-level steering were explored and can partially reduce EM.
Topics
Related items
- CriticalHermes Agent - PKCE Session Takeover via Redirect-URI Parser ConfusionSimilar attack · Tenable Research Advisories
- LowLost in the comments: Social context as a single‐pass jailbreak and defense on agentic platformsSimilar attack · OpenAlex (peer-reviewed AI security)
- MediumGHSA-hmq2-7hp6-7crh: Banks: User-controlled prompt input can be parsed as privileged chat messagesSimilar attack · GitHub Advisory Database
- HighGHSA-6wjp-v33h-5cvq: PraisonAI: AgentOS defaults to network-exposed no-auth mode, allowing unauthenticated agent invocation and instruction disclosureSimilar attack · GitHub Advisory Database
- HighCVE-2026-101998: Docker Sandboxes fail open when masking credentials in proxy responsesSimilar attack · NVD/CVE Database