InfoResearchPreprint
Detect and Suppress: A Mechanistic Defense against Adversarial Patches in VLA Models
- Published
- Record updated
Summary
Researchers analyze Vision-Language-Action (VLA) models with a sparse autoencoder (SAE) and find an internal feature whose activation strongly correlates with adversarial patches. They suppress this feature at inference time only when a linear probe detects an attack, which improves robustness without fine-tuning the VLA. On LIBERO-10, conditional intervention raises success rate under intermittent attacks, while continuous intervention substantially degrades policy performance.
Mitigation
Suppress the identified SAE feature at inference time, applying the intervention only when a linear probe detects an attack. Avoid continuous application, which substantially degrades policy performance.
Topics
Related items
- InfoDoes Target Alignment Mean Target Recovery? An Evidence-Ladder Study of Adversarial Claims on Contrastive EncodersSimilar attack · Arxiv (cs.RO + cs.CV security)
- InfoBRANCH: Bypassing Multi-Scanner AI GuardrailsSimilar attack · Arxiv (cs.CR + cs.CL + cs.LG)
- InfoDetecting Adversarial Images through Response Profiles of Vision-Language ModelsSimilar attack · Arxiv (cs.RO + cs.CV security)
- InfoGraphRectify: Graph-Based Transfer of Adversarial Example Detectors Across Neural NetworksSimilar attack · Arxiv (cs.RO + cs.CV security)
- InfoVCR-Bench: A Modular Open-Source Benchmark for Video Classification RobustnessSimilar attack · Arxiv (cs.RO + cs.CV security)