Beyond Single-Pair Attacks: Disrupting Vision-Language Pre-Training Models With Dual-Semantic Frequency Stealth
Summary
Researchers have discovered a new attack method called DSFG-Attack that can fool Vision-Language Pre-training models (AI systems trained to understand both images and text together) by creating adversarial examples (slightly altered inputs designed to trick AI). The attack works by injecting conflicting information between images and text, and hiding the changes in high-frequency image details (fine textures), making the attack harder to detect and more effective at transferring between different AI systems, including advanced models like GPT-4o.
Classification
Affected Vendors
Related Issues
Original source: http://ieeexplore.ieee.org/document/11609281
First tracked: September 3, 2026 at 08:02 PM
Classified by LLM (prompt v3) · confidence: 92%