InfoResearchPreprint
Backdoor as Probe: Test-Time Adversarial Defense for CLIP
- Published
- Record updated
Summary
Researchers propose Backdoor as Probe (BaP), a test-time adversarial defense for CLIP that repurposes backdoor trigger-to-target mechanisms. A defender-controlled probe is implanted via a closed-form edit to a selected MLP layer, responding weakly to clean inputs and strongly to adversarial activation shifts. Across 16 benchmarks, BaP raises average robust accuracy from 1.0% to 52.3% while keeping clean accuracy, with up to a 5.7x inference speedup.
Mitigation
BaP is the mitigation: a test-time defense that detects adversarial inputs via the probe's target-direction response, then selectively rectifies them by optimizing a small perturbation that steers representations toward the clean subspace.
Related items
- InfoDoes Target Alignment Mean Target Recovery? An Evidence-Ladder Study of Adversarial Claims on Contrastive EncodersSimilar attack · Arxiv (cs.RO + cs.CV security)
- InfoBRANCH: Bypassing Multi-Scanner AI GuardrailsSimilar attack · Arxiv (cs.CR + cs.CL + cs.LG)
- InfoDetecting Adversarial Images through Response Profiles of Vision-Language ModelsSimilar attack · Arxiv (cs.RO + cs.CV security)
- InfoGraphRectify: Graph-Based Transfer of Adversarial Example Detectors Across Neural NetworksSimilar attack · Arxiv (cs.RO + cs.CV security)
- InfoVCR-Bench: A Modular Open-Source Benchmark for Video Classification RobustnessSimilar attack · Arxiv (cs.RO + cs.CV security)