Skip to content
InfoResearchPreprintLLM-specific

One Attack to Fool Them All: Highly Transferable Black-Box Adversarial Attacks on Frontier MLLMs

Published
Record updated
View JSON

Summary

Researchers ask whether a single adversarial image can consistently mislead diverse frontier MLLMs in black-box settings. They present O-Attack, a transfer-based black-box framework that exploits cross-modally aligned semantic representations in surrogate models, and report attack success rates rising on GPT-5.4 (29.1% to 77.2%), Claude-4.6 (42.8% to 81.6%), and Gemini-3.1 (38.2% to 80.9%). Across 24 MLLMs, O-Attack outperforms six state-of-the-art methods in black-box transferability.