Skip to content
InfoResearchPreprintLLM-specific

Visual-Invariance-Augmented Feature Optimal Alignment for Transferable Adversarial Attacks against Closed-Source MLLMs

Published
Record updated
View JSON

Summary

Researchers study how to make adversarial images transfer from open-source surrogate multimodal LLMs to closed-source MLLMs in black-box settings. They propose IAU-FOA, which aligns adversarial and target images at both global and patch-cluster levels using confidence-adaptive unbalanced optimal transport, plus visual-invariance augmentation that simulates exposure, contrast, illumination and color-temperature changes. The authors report that it consistently outperforms state-of-the-art transferable attack methods across open-source and closed-source MLLMs.