Skip to content
InfoResearchPreprint

Does Target Alignment Mean Target Recovery? An Evidence-Ladder Study of Adversarial Claims on Contrastive Encoders

Published
Record updated
View JSON

Summary

This research asks whether the victim-space target alignment score used in adversarial attacks on vision-language models predicts recovery of the target by an independent model. Across six contrastive encoders under a matched attack at three perturbation budgets, robustly trained encoders (FARE, TeCoA, PMG, TRADES) transfer substantially more independent evidence than vanilla CLIP or SigLIP, but no cell reaches the blend-derived reference level, so practical recovery is left undecided. The authors conclude the score is informative only within a fixed robust encoder and offer a reporting protocol in its place.