Skip to content
InfoResearchPreprint

Detect and Suppress: A Mechanistic Defense against Adversarial Patches in VLA Models

Published
Record updated
View JSON

Summary

Researchers analyze Vision-Language-Action (VLA) models with a sparse autoencoder (SAE) and find an internal feature whose activation strongly correlates with adversarial patches. They suppress this feature at inference time only when a linear probe detects an attack, which improves robustness without fine-tuning the VLA. On LIBERO-10, conditional intervention raises success rate under intermittent attacks, while continuous intervention substantially degrades policy performance.

Mitigation

Suppress the identified SAE feature at inference time, applying the intervention only when a linear probe detects an attack. Avoid continuous application, which substantially degrades policy performance.