InfoResearchPreprint
Selective Channel Restoration for Backdoored Vision-Language Models
- Published
- Record updated
Summary
Researchers propose Perturb-Select-Restore (PSR), a post-training defense against backdoors in vision-language models (VLMs) implanted through poisoned fine-tuning data. PSR performs sparse updates to the projection interface and adds no computation at inference time. The authors report that backdoored VLM projectors are more sensitive to bounded perturbations than clean ones, a property they call projection fragility, and PSR reduces attack success rates to near zero while preserving clean-task performance.
Mitigation
PSR, a post-training defense that identifies the output channels most sensitive to perturbations in each projection layer of a backdoored VLM and restores their parameters to the corresponding pretrained values.
Topics
Related items
- LowAnthropic Cuts Live Internet Access for Internal AI Tests After Claude Exploits Injection FlawsSimilar attack · The Hacker News
- CriticalCVE-2026-108263: Astron Agent code-node execution as root through workflow run endpointsSimilar attack · NVD/CVE Database
- MediumHackers abuse Google Ads, Bing redirects to push Claude ClickFix attacksSimilar attack · BleepingComputer
- CriticalHermes Agent - PKCE Session Takeover via Redirect-URI Parser ConfusionSimilar attack · Tenable Research Advisories
- LowSocial Engineering AI Agents: The New BEC for 2026Similar attack · Dark Reading