InfoResearchPreprint
Detecting Adversarial Images through Response Profiles of Vision-Language Models
- Published
- Record updated
Summary
The paper proposes a detector that identifies adversarial images for frozen vision-language models by profiling how an image responds to a set of general semantic prompts. The profile combines category-level statistics, prompt relationships, deviations from clean reference distributions, and stability under weak image transformations, and a lightweight classifier labels each input while the VLM stays fixed. Evaluated across multiple datasets, CLIP-style backbones and several attack families, the detector discriminates strongly in attack-specific settings and retains substantial performance on unseen attacks.
Topics
Related items
- InfoDoes Target Alignment Mean Target Recovery? An Evidence-Ladder Study of Adversarial Claims on Contrastive EncodersSimilar attack · Arxiv (cs.RO + cs.CV security)
- InfoBRANCH: Bypassing Multi-Scanner AI GuardrailsSimilar attack · Arxiv (cs.CR + cs.CL + cs.LG)
- InfoGraphRectify: Graph-Based Transfer of Adversarial Example Detectors Across Neural NetworksSimilar attack · Arxiv (cs.RO + cs.CV security)
- InfoVCR-Bench: A Modular Open-Source Benchmark for Video Classification RobustnessSimilar attack · Arxiv (cs.RO + cs.CV security)
- InfoTransferable Spatial Temporal Coherence Adversarial Attack on Black-Box Vision Language Models for Autonomous DrivingSimilar attack · Arxiv (cs.RO + cs.CV security)