InfoResearchPreprintLLM-specific
Corrupted but Correct: Why Vision-Language Models Lie to Themselves Internally
- Published
- Record updated
Summary
Researchers report that a targeted adversarial perturbation can drive a vision-language model's teacher-forced training loss for a fixed target caption near zero, while free generation on Qwen2.5-VL-7B-Instruct still yields the correct description. Using 200 held-out COCO images and a two-stage PGD attack, they localise the gap to one autoregressive step and across the 28 LLM decoder layers, finding that the language decoder, not the visual encoder, differentially arbitrates which images are corrupted (linear probe AUC=0.858, flagged for circularity). The authors argue that adversarial robustness in autoregressive VLMs depends mainly on the language decoder's prior.
Topics
Related items
- InfoDoes Target Alignment Mean Target Recovery? An Evidence-Ladder Study of Adversarial Claims on Contrastive EncodersSimilar attack · Arxiv (cs.RO + cs.CV security)
- InfoBRANCH: Bypassing Multi-Scanner AI GuardrailsSimilar attack · Arxiv (cs.CR + cs.CL + cs.LG)
- InfoDetecting Adversarial Images through Response Profiles of Vision-Language ModelsSimilar attack · Arxiv (cs.RO + cs.CV security)
- InfoGraphRectify: Graph-Based Transfer of Adversarial Example Detectors Across Neural NetworksSimilar attack · Arxiv (cs.RO + cs.CV security)
- InfoVCR-Bench: A Modular Open-Source Benchmark for Video Classification RobustnessSimilar attack · Arxiv (cs.RO + cs.CV security)