Skip to content
InfoResearchPreprintLLM-specific

It Takes Little to Rewrite Perception: Targeted Semantic Substitution in Vision-Language Models at $ε\leq 4/255$

Published
Record updated
View JSON

Summary

Researchers show that targeted semantic substitution can make vision-language models (VLMs) perceive a different target than the source image within an adversarial perturbation budget of ε ≤ 4/255, a range where prior representation-alignment attacks achieved limited success. Under a white-box threat model, the source image's token streams are aligned with the target's in the victim VLM's post-merger token space. Under a strict success criterion, complete replacement reaches 38% on images at ε = 4/255 and 35.9% on video at ε = 1/255.