InfoResearchPreprint
Does Target Alignment Mean Target Recovery? An Evidence-Ladder Study of Adversarial Claims on Contrastive Encoders
- Published
- Record updated
Summary
This research asks whether the victim-space target alignment score used in adversarial attacks on vision-language models predicts recovery of the target by an independent model. Across six contrastive encoders under a matched attack at three perturbation budgets, robustly trained encoders (FARE, TeCoA, PMG, TRADES) transfer substantially more independent evidence than vanilla CLIP or SigLIP, but no cell reaches the blend-derived reference level, so practical recovery is left undecided. The authors conclude the score is informative only within a fixed robust encoder and offer a reporting protocol in its place.
Related items
- InfoBRANCH: Bypassing Multi-Scanner AI GuardrailsSimilar attack · Arxiv (cs.CR + cs.CL + cs.LG)
- InfoDetecting Adversarial Images through Response Profiles of Vision-Language ModelsSimilar attack · Arxiv (cs.RO + cs.CV security)
- InfoGraphRectify: Graph-Based Transfer of Adversarial Example Detectors Across Neural NetworksSimilar attack · Arxiv (cs.RO + cs.CV security)
- InfoVCR-Bench: A Modular Open-Source Benchmark for Video Classification RobustnessSimilar attack · Arxiv (cs.RO + cs.CV security)
- InfoTransferable Spatial Temporal Coherence Adversarial Attack on Black-Box Vision Language Models for Autonomous DrivingSimilar attack · Arxiv (cs.RO + cs.CV security)