Purifying Backdoored Large Vision-Language Models by Removing Hijacked Directions
- Published
- Record updated
Summary
Researchers propose OrthoPurify, a method that removes backdoors from large vision-language models (LVLMs) by applying a one-step orthogonal projection to the backdoored weight updates. The method identifies the backdoor as "direction hijacking," where a small number of weight update directions are diverted from task adaptation to encode a backdoor shortcut. A pseudo-benign reference model, fine-tuned on a small set of clean samples, is used to isolate these hijacked directions. The authors report that OrthoPurify reduces attack success rate to near zero while preserving original performance, without retraining or inference-time overhead.
Mitigation
OrthoPurify: a one-step orthogonal projection on the weight update that removes hijacked directions, using a pseudo-benign reference model obtained by fine-tuning the pretrained weights on a small set of clean samples. Code is available at https://github.com/womeimingzi/OrthoPurify.
Topics
Related items
- LowAnthropic Cuts Live Internet Access for Internal AI Tests After Claude Exploits Injection FlawsSimilar attack · The Hacker News
- CriticalCVE-2026-108263: Astron Agent code-node execution as root through workflow run endpointsSimilar attack · NVD/CVE Database
- MediumHackers abuse Google Ads, Bing redirects to push Claude ClickFix attacksSimilar attack · BleepingComputer
- CriticalHermes Agent - PKCE Session Takeover via Redirect-URI Parser ConfusionSimilar attack · Tenable Research Advisories
- LowSocial Engineering AI Agents: The New BEC for 2026Similar attack · Dark Reading