InfoResearchPreprintLLM-specific
It Takes Little to Rewrite Perception: Targeted Semantic Substitution in Vision-Language Models at $ε\leq 4/255$
- Published
- Record updated
Summary
Researchers show that targeted semantic substitution can make vision-language models (VLMs) perceive a different target than the source image within an adversarial perturbation budget of ε ≤ 4/255, a range where prior representation-alignment attacks achieved limited success. Under a white-box threat model, the source image's token streams are aligned with the target's in the victim VLM's post-merger token space. Under a strict success criterion, complete replacement reaches 38% on images at ε = 4/255 and 35.9% on video at ε = 1/255.
Topics
Related items
- InfoDoes Target Alignment Mean Target Recovery? An Evidence-Ladder Study of Adversarial Claims on Contrastive EncodersSimilar attack · Arxiv (cs.RO + cs.CV security)
- InfoBRANCH: Bypassing Multi-Scanner AI GuardrailsSimilar attack · Arxiv (cs.CR + cs.CL + cs.LG)
- InfoDetecting Adversarial Images through Response Profiles of Vision-Language ModelsSimilar attack · Arxiv (cs.RO + cs.CV security)
- InfoGraphRectify: Graph-Based Transfer of Adversarial Example Detectors Across Neural NetworksSimilar attack · Arxiv (cs.RO + cs.CV security)
- InfoVCR-Bench: A Modular Open-Source Benchmark for Video Classification RobustnessSimilar attack · Arxiv (cs.RO + cs.CV security)