Multimodal models
Models that read or produce images, audio or video, and attacks that hide inside those inputs.
- All items
- 19
- Last 90 days
- 10
- Change
- +100%vs 5 before
Items per month
| Month | Items |
|---|---|
| May 2025 | 0 |
| Jun 2025 | 0 |
| Jul 2025 | 0 |
| Aug 2025 | 0 |
| Sep 2025 | 0 |
| Oct 2025 | 0 |
| Nov 2025 | 0 |
| Dec 2025 | 0 |
| Jan 2026 | 1 |
| Feb 2026 | 1 |
| Mar 2026 | 2 |
| Apr 2026 | 2 |
| May 2026 | 1 |
| Jun 2026 | 2 |
| Jul 2026 | 1 |
| Aug 2026 | 0 |
| Sep 2026 | 4 |
| Oct 2026 | 5 |
17 items
Detecting Adversarial Images through Response Profiles of Vision-Language Models
Oct 7, 2026InfoResearchPreprintSecurityResearchThe paper proposes a detector that identifies adversarial images for frozen vision-language models by profiling how an image responds to a set of general semantic prompts. The profile combines category-level statistics, prompt relationships, deviations from clean reference distributions, and stability under weak image transformations, and a lightweight classifier labels each input while the VLM stays fixed. Evaluated across multiple datasets, CLIP-style backbones and several attack families, the detector discriminates strongly in attack-specific settings and retains substantial performance on unseen attacks.
Arxiv (cs.RO + cs.CV security)Purifying Backdoored Large Vision-Language Models by Removing Hijacked Directions
Oct 7, 2026InfoResearchPreprintSecurityResearchResearchers propose OrthoPurify, a method that removes backdoors from large vision-language models (LVLMs) by applying a one-step orthogonal projection to the backdoored weight updates. The method identifies the backdoor as "direction hijacking," where a small number of weight update directions are diverted from task adaptation to encode a backdoor shortcut. A pseudo-benign reference model, fine-tuned on a small set of clean samples, is used to isolate these hijacked directions. The authors report that OrthoPurify reduces attack success rate to near zero while preserving original performance, without retraining or inference-time overhead.
Fix: OrthoPurify: a one-step orthogonal projection on the weight update that removes hijacked directions, using a pseudo-benign reference model obtained by fine-tuning the pretrained weights on a small set of clean samples. Code is available at https://github.com/womeimingzi/OrthoPurify.
Arxiv (cs.RO + cs.CV security)Vulnerabilities and Defenses in Audio-Visual Attacks: A Survey From Audio to Multimodal Models
Oct 4, 2026InfoResearchPeer-reviewedSecurityResearchThis survey reviews attacks on audio-visual multimodal large language models (MLLMs), covering adversarial, backdoor and jailbreak attacks. It notes that researchers often fine-tune public open-source MLLMs, which introduces security risks, and that existing surveys address only specific attack types. The paper also reviews attacks against the latest audio-visual MLLMs and outlines challenges and trends for future research on attacks and defenses. Published in ACM Computing Surveys on 2026-10-05.
OpenAlex (peer-reviewed AI security)Image-embedded prompt injection vulnerability of vision-language models in dental radiology: a cross-vendor attack–defense evaluation
Oct 2, 2026InfoResearchPeer-reviewedSecurityResearchResearchers evaluated image-embedded prompt injection, where adversarial text is rendered into medical image pixels, against four vision-language models (GPT-4o, Gemini 2.5 Flash, Claude Sonnet 4.5, MedGemma 4B) using 270 dental panoramic radiographs from the DenTeX dataset. All four models were vulnerable, with paired attack success rates up to 62.6% (95% CI: 58.5–66.7%) for GPT-4o. Among five benchmarked defenses, OCR-based text sanitization achieved the strongest reduction (pooled ASR: 0.2%), while the provenance-aware ProvDent defense escalates suspicious cases for human review and kept clean-image F1 within 0.6 percentage points of baseline.
Fix: OCR-based text sanitization achieved the strongest attack reduction (pooled ASR: 0.2%). The provenance-aware ProvDent defense provides a complementary fail-open mechanism that escalates suspicious cases for human review.
OpenAlex (peer-reviewed AI security)Corrupted but Correct: Why Vision-Language Models Lie to Themselves Internally
Oct 2, 2026InfoResearchPreprintSecurityResearchResearchers report that a targeted adversarial perturbation can drive a vision-language model's teacher-forced training loss for a fixed target caption near zero, while free generation on Qwen2.5-VL-7B-Instruct still yields the correct description. Using 200 held-out COCO images and a two-stage PGD attack, they localise the gap to one autoregressive step and across the 28 LLM decoder layers, finding that the language decoder, not the visual encoder, differentially arbitrates which images are corrupted (linear probe AUC=0.858, flagged for circularity). The authors argue that adversarial robustness in autoregressive VLMs depends mainly on the language decoder's prior.
Arxiv (cs.RO + cs.CV security)It Takes Little to Rewrite Perception: Targeted Semantic Substitution in Vision-Language Models at $ε\leq 4/255$
Sep 29, 2026InfoResearchPreprintSecurityResearchResearchers show that targeted semantic substitution can make vision-language models (VLMs) perceive a different target than the source image within an adversarial perturbation budget of ε ≤ 4/255, a range where prior representation-alignment attacks achieved limited success. Under a white-box threat model, the source image's token streams are aligned with the target's in the victim VLM's post-merger token space. Under a strict success criterion, complete replacement reaches 38% on images at ε = 4/255 and 35.9% on video at ε = 1/255.
Arxiv (cs.RO + cs.CV security)Selective Channel Restoration for Backdoored Vision-Language Models
Sep 29, 2026InfoResearchPreprintSecurityResearchResearchers propose Perturb-Select-Restore (PSR), a post-training defense against backdoors in vision-language models (VLMs) implanted through poisoned fine-tuning data. PSR performs sparse updates to the projection interface and adds no computation at inference time. The authors report that backdoored VLM projectors are more sensitive to bounded perturbations than clean ones, a property they call projection fragility, and PSR reduces attack success rates to near zero while preserving clean-task performance.
Fix: PSR, a post-training defense that identifies the output channels most sensitive to perturbations in each projection layer of a backdoored VLM and restores their parameters to the corresponding pretrained values.
Arxiv (cs.RO + cs.CV security)Narrow Multimodal Fine-Tuning Can Induce Emergent Misalignment
Sep 28, 2026InfoResearchPreprintSafetyResearchResearchers study emergent misalignment (EM) in vision-language models, where fine-tuning on narrow multimodal tasks causes broadly harmful behavior. Across fifteen commercial and open-source models, narrow multimodal fine-tuning induced misaligned behavior that transferred to unrelated tasks, including visual factual dishonesty, unsafe image generation, visual jailbreak vulnerability and risky agentic actions. The authors report that EM arises under both supervised fine-tuning and preference optimization and is sensitive to training-evaluation modality alignment.
Fix: Prompt inoculation, benign continued training, and activation-level steering were explored and can partially reduce EM.
Arxiv (cs.RO + cs.CV security)Empirical Analysis of Goal Hijacking in Large Vision-Language Models via Visual Prompt Injection
Sep 27, 2026InfoResearchPeer-reviewedSecurityResearchResearchers study visual prompt injection (VPI), where instructions embedded in input images are followed by large vision-language models (LVLMs). They propose "goal hijacking via visual prompt injection" (GHVPI), which redirects an LVLM from its original task to an attacker-specified one. Their quantitative analysis reports an attack success rate of 15.8% against GPT-4V, and they find GHVPI success depends on the character recognition and instruction-following capabilities of LVLMs.
OpenAlex (peer-reviewed AI security)Enhancing Targeted Adversarial Attacks on Large Vision-Language Models via Intermediate Projector
Jun 10, 2026InfoResearchPeer-reviewedSecurityResearchResearchers show that black-box targeted attacks on Large Vision-Language Models can be made more precise by working through the projector, a semantic bridge between vision and language. They propose Intermediate Projector Guided Attack (IPGA), which aligns Q-Former query outputs with a target and transfers across models, and IPGA with Residual Query Alignment (IPGA-R), which also preserves non-target content for fine-grained edits. The authors report that IPGA beats baselines on global attacks, IPGA-R wins on fine-grained attacks, and the method transfers to Google Gemini and OpenAI GPT.
IEEE Xplore (Security & AI Journals)PVLM: Parsing-Aware Vision-Language Model With Dynamic Contrastive Learning for Zero-Shot Deepfake Attribution
May 14, 2026InfoResearchPeer-reviewedSecurityResearchPVLM is a parsing-aware vision-language model with dynamic contrastive learning for zero-shot deepfake attribution, meaning it traces forged faces to generators not seen in training, including diffusion models. The authors build a fine-grained ZS-DFA benchmark and use face parsing to exploit differences in how GAN and diffusion generators preserve source facial attributes. They report that the model exceeds the state of the art on the ZS-DFA benchmark across various protocol evaluations.
IEEE Xplore (Security & AI Journals)Privacy-preserving for user-uploaded images and text in Vision-Language Models
Apr 28, 2026InfoResearchPeer-reviewedPrivacyResearchElsevier Security JournalsVLBiasBench: A Comprehensive Benchmark for Evaluating Bias in Large Vision-Language Model
Apr 14, 2026InfoResearchPeer-reviewedResearchSafetyVLBiasBench is a benchmark for measuring social bias in Large Vision-Language Models (LVLMs). It covers nine bias categories plus two intersectional ones (race × gender and race × social economic status). The dataset has 128,342 samples built from 46,848 images generated with Stable Diffusion XL, and the authors evaluated 15 open-source models and two closed-source models.
IEEE Xplore (Security & AI Journals)SMInject: Specious Malignant Injection Attacks With Semantically-Enhanced Tokens in Cross-Modal Retrieval
Mar 13, 2026InfoResearchPeer-reviewedSecurityResearchSMInject is a new injection attack framework against pre-trained multimodal models used in cross-modal retrieval. It generates deceptive injections that combine concepts through causal correlation across modalities, and aligns them in the encoded embedding space to boost effectiveness. On representative multimodal models, it reportedly achieves over 14% higher attack success rate and 6% higher Hit@5 than state-of-the-art methods while preserving overall model utility.
IEEE Xplore (Security & AI Journals)Are Large Vision-Language Models Robust to Adversarial Visual Transformations?
Mar 5, 2026InfoResearchPeer-reviewedSecurityResearchResearchers assess how robust large vision-language models (LVLMs) are to adversarial visual transformations, a simpler attack style than optimizing perturbations or manipulating prompts. They test LVLM resilience across all possible transformation operations and find that combining the most harmful transformations yields more effective attacks. They also introduce adversarial learning of visual transformations, which applies malicious transformations to raw images via gradient approximation to improve attack effectiveness and imperceptibility.
IEEE Xplore (Security & AI Journals)Evaluating and Mitigating Relationship Hallucinations in Large Vision-Language Models
Feb 3, 2026InfoResearchPeer-reviewedResearchSafetyR-Bench is a new benchmark for measuring hallucinations about relationships between objects in Large Vision-Language Models, using image-level questions on whether a relationship exists and instance-level questions on local visual comprehension. The authors trace these hallucinations to relationship-relationship, subject-relationship and relationship-object co-occurrences, worsened by long-tail distributions in visual datasets. They report that region-level image-text alignment reduces them and propose a baseline, Region-Aware Alignment Mitigation (RA$^{2}$2M), that directs model attention to relevant regions.
Fix: Region-level image-text alignment helps mitigate relationship hallucinations; the authors propose RA$^{2}$2M (Region-Aware Alignment Mitigation) as a new baseline that enhances model attention to relevant regions.
IEEE Xplore (Security & AI Journals)NAP-Tuning: Neural Augmented Prompt Tuning for Adversarially Robust Vision-Language Models
Jan 30, 2026InfoResearchPeer-reviewedSecurityResearchNAP-Tuning extends Adversarial Prompt Tuning (AdvPT) for vision-language models such as CLIP by adding a multi-modal, multi-layer prompting framework. Its core is a Neural Augmentor that uses TokenRefiners, lightweight modules that reconstruct purified features through residual connections to correct adversarial distortions in feature space. Under the AutoAttack benchmark it outperforms the strongest baselines by 32.3% on ViT-B16 and 31.3% on ViT-B32 while keeping competitive clean accuracy.
IEEE Xplore (Security & AI Journals)
Topic added 2026-10-09. An item belongs to this topic when its title matches one of the topic's patterns or its summary mentions the topic at least twice. Report a wrong match with the feedback button on the item.