Skip to content
InfoResearchPreprint

Purifying Backdoored Large Vision-Language Models by Removing Hijacked Directions

Published
Record updated
View JSON

Summary

Researchers propose OrthoPurify, a method that removes backdoors from large vision-language models (LVLMs) by applying a one-step orthogonal projection to the backdoored weight updates. The method identifies the backdoor as "direction hijacking," where a small number of weight update directions are diverted from task adaptation to encode a backdoor shortcut. A pseudo-benign reference model, fine-tuned on a small set of clean samples, is used to isolate these hijacked directions. The authors report that OrthoPurify reduces attack success rate to near zero while preserving original performance, without retraining or inference-time overhead.

Mitigation

OrthoPurify: a one-step orthogonal projection on the weight update that removes hijacked directions, using a pseudo-benign reference model obtained by fine-tuning the pretrained weights on a small set of clean samples. Code is available at https://github.com/womeimingzi/OrthoPurify.