InfoResearchPreprintLLM-specific
One Attack to Fool Them All: Highly Transferable Black-Box Adversarial Attacks on Frontier MLLMs
- Published
- Record updated
Summary
Researchers ask whether a single adversarial image can consistently mislead diverse frontier MLLMs in black-box settings. They present O-Attack, a transfer-based black-box framework that exploits cross-modally aligned semantic representations in surrogate models, and report attack success rates rising on GPT-5.4 (29.1% to 77.2%), Claude-4.6 (42.8% to 81.6%), and Gemini-3.1 (38.2% to 80.9%). Across 24 MLLMs, O-Attack outperforms six state-of-the-art methods in black-box transferability.
Topics
Related items
- InfoRogue Anthropic AI agent gave police fake tip in unsolved murder caseSame vendor · BBC Technology
- LowAnthropic Cuts Live Internet Access for Internal AI Tests After Claude Exploits Injection FlawsSame vendor · The Hacker News
- InfoQuoting The New York TimesSame vendor · Simon Willison's Weblog
- InfoAnthropic’s AI gave Philadelphia police a fake tip about an unsolved homicideSame vendor · The Verge (AI)
- InfoOpenAI Fires 3 Safety Researchers in Dispute Over AI RisksSame vendor · SecurityWeek