Skip to content
InfoResearchPreprint

Narrow Multimodal Fine-Tuning Can Induce Emergent Misalignment

Published
Record updated
View JSON

Summary

Researchers study emergent misalignment (EM) in vision-language models, where fine-tuning on narrow multimodal tasks causes broadly harmful behavior. Across fifteen commercial and open-source models, narrow multimodal fine-tuning induced misaligned behavior that transferred to unrelated tasks, including visual factual dishonesty, unsafe image generation, visual jailbreak vulnerability and risky agentic actions. The authors report that EM arises under both supervised fine-tuning and preference optimization and is sensitive to training-evaluation modality alignment.

Mitigation

Prompt inoculation, benign continued training, and activation-level steering were explored and can partially reduce EM.