{"data":[{"id":"303b8d4b-1efb-4ca7-8a69-0cd5f33fccc3","title":"Detecting Adversarial Images through Response Profiles of Vision-Language Models","headline":null,"summary":"The paper proposes a detector that identifies adversarial images for frozen vision-language models by profiling how an image responds to a set of general semantic prompts. The profile combines category-level statistics, prompt relationships, deviations from clean reference distributions, and stability under weak image transformations, and a lightweight classifier labels each input while the VLM stays fixed. Evaluated across multiple datasets, CLIP-style backbones and several attack families, the detector discriminates strongly in attack-specific settings and retains substantial performance on unseen attacks.","sourceUrl":"https://arxiv.org/abs/2610.10436v1","publishedAt":"2026-10-07T17:13:12.000Z","severity":"info","cvssSeverity":null,"cvssScore":null,"labels":["security","research"],"issueType":"research","cveId":null,"cweIds":null,"affectedPackages":null,"affectedVendors":[],"affectedVendorsRaw":["CLIP-style vision-language models"],"classifierModel":"claude-haiku-5-5","classifierPromptVersion":"v4","solution":"N/A -- no mitigation discussed in source.","attackType":["model_evasion"],"cvssVector":null,"attackVector":null,"attackComplexity":null,"privilegesRequired":null,"userInteraction":null,"exploitMaturity":null,"epssScore":null,"epssCheckedAt":null,"kevDateAdded":null,"advisoryAliases":null,"affectedPackagesSource":null,"patchAvailable":null,"disclosureDate":"2026-10-07T17:13:12.000Z","capecIds":null,"crossRefCount":0,"attackSophistication":"moderate","impactType":["integrity"],"aiComponentTargeted":"model","llmSpecific":false,"classifierConfidence":0.93,"researchCategory":"preprint","atlasIds":null},{"id":"75f4edc4-5482-4d8c-919e-81bc8d4fd984","title":"Purifying Backdoored Large Vision-Language Models by Removing Hijacked Directions","headline":null,"summary":"Researchers propose OrthoPurify, a method that removes backdoors from large vision-language models (LVLMs) by applying a one-step orthogonal projection to the backdoored weight updates. The method identifies the backdoor as \"direction hijacking,\" where a small number of weight update directions are diverted from task adaptation to encode a backdoor shortcut. A pseudo-benign reference model, fine-tuned on a small set of clean samples, is used to isolate these hijacked directions. The authors report that OrthoPurify reduces attack success rate to near zero while preserving original performance, without retraining or inference-time overhead.","sourceUrl":"https://arxiv.org/abs/2610.09941v1","publishedAt":"2026-10-07T12:25:47.000Z","severity":"info","cvssSeverity":null,"cvssScore":null,"labels":["security","research"],"issueType":"research","cveId":null,"cweIds":null,"affectedPackages":null,"affectedVendors":[],"affectedVendorsRaw":["LVLMs"],"classifierModel":"claude-haiku-5-5","classifierPromptVersion":"v4","solution":"OrthoPurify: a one-step orthogonal projection on the weight update that removes hijacked directions, using a pseudo-benign reference model obtained by fine-tuning the pretrained weights on a small set of clean samples. Code is available at https://github.com/womeimingzi/OrthoPurify.","attackType":["other"],"cvssVector":null,"attackVector":null,"attackComplexity":null,"privilegesRequired":null,"userInteraction":null,"exploitMaturity":null,"epssScore":null,"epssCheckedAt":null,"kevDateAdded":null,"advisoryAliases":null,"affectedPackagesSource":null,"patchAvailable":null,"disclosureDate":"2026-10-07T12:25:47.000Z","capecIds":null,"crossRefCount":0,"attackSophistication":"advanced","impactType":["integrity","safety"],"aiComponentTargeted":"model","llmSpecific":false,"classifierConfidence":0.95,"researchCategory":"preprint","atlasIds":null},{"id":"d3633832-d3ca-4aa4-9828-4e726daecda5","title":"Vulnerabilities and Defenses in Audio-Visual Attacks: A Survey From Audio to Multimodal Models","headline":null,"summary":"This survey reviews attacks on audio-visual multimodal large language models (MLLMs), covering adversarial, backdoor and jailbreak attacks. It notes that researchers often fine-tune public open-source MLLMs, which introduces security risks, and that existing surveys address only specific attack types. The paper also reviews attacks against the latest audio-visual MLLMs and outlines challenges and trends for future research on attacks and defenses. Published in ACM Computing Surveys on 2026-10-05.","sourceUrl":"https://doi.org/10.1145/3857219","publishedAt":"2026-10-05T00:00:00.000Z","severity":"info","cvssSeverity":null,"cvssScore":null,"labels":["security","research"],"issueType":"research","cveId":null,"cweIds":null,"affectedPackages":null,"affectedVendors":[],"affectedVendorsRaw":["multimodal large language models","audio-visual MLLMs"],"classifierModel":"claude-haiku-5-5","classifierPromptVersion":"v4","solution":"N/A -- no mitigation discussed in source.","attackType":["jailbreak","model_evasion"],"cvssVector":null,"attackVector":null,"attackComplexity":null,"privilegesRequired":null,"userInteraction":null,"exploitMaturity":null,"epssScore":null,"epssCheckedAt":null,"kevDateAdded":null,"advisoryAliases":null,"affectedPackagesSource":null,"patchAvailable":null,"disclosureDate":"2026-10-05T00:00:00.000Z","capecIds":null,"crossRefCount":0,"attackSophistication":"moderate","impactType":["integrity","safety"],"aiComponentTargeted":"model","llmSpecific":true,"classifierConfidence":0.9,"researchCategory":"peer_reviewed","atlasIds":null},{"id":"589321c7-340f-46f1-a93f-d76f9f2d0776","title":"Image-embedded prompt injection vulnerability of vision-language models in dental radiology: a cross-vendor attack–defense evaluation","headline":null,"summary":"Researchers evaluated image-embedded prompt injection, where adversarial text is rendered into medical image pixels, against four vision-language models (GPT-4o, Gemini 2.5 Flash, Claude Sonnet 4.5, MedGemma 4B) using 270 dental panoramic radiographs from the DenTeX dataset. All four models were vulnerable, with paired attack success rates up to 62.6% (95% CI: 58.5–66.7%) for GPT-4o. Among five benchmarked defenses, OCR-based text sanitization achieved the strongest reduction (pooled ASR: 0.2%), while the provenance-aware ProvDent defense escalates suspicious cases for human review and kept clean-image F1 within 0.6 percentage points of baseline.","sourceUrl":"https://doi.org/10.1038/s41598-026-74077-3","publishedAt":"2026-10-03T00:00:00.000Z","severity":"info","cvssSeverity":null,"cvssScore":null,"labels":["security","research"],"issueType":"research","cveId":null,"cweIds":null,"affectedPackages":null,"affectedVendors":["OpenAI","Google","Anthropic"],"affectedVendorsRaw":["GPT-4o","Gemini 2.5 Flash","Claude Sonnet 4.5","MedGemma 4B"],"classifierModel":"claude-haiku-5-5","classifierPromptVersion":"v4","solution":"OCR-based text sanitization achieved the strongest attack reduction (pooled ASR: 0.2%). The provenance-aware ProvDent defense provides a complementary fail-open mechanism that escalates suspicious cases for human review.","attackType":["prompt_injection"],"cvssVector":null,"attackVector":null,"attackComplexity":null,"privilegesRequired":null,"userInteraction":null,"exploitMaturity":null,"epssScore":null,"epssCheckedAt":null,"kevDateAdded":null,"advisoryAliases":null,"affectedPackagesSource":null,"patchAvailable":null,"disclosureDate":"2026-10-03T00:00:00.000Z","capecIds":null,"crossRefCount":0,"attackSophistication":"moderate","impactType":["integrity","safety"],"aiComponentTargeted":"model","llmSpecific":true,"classifierConfidence":0.95,"researchCategory":"peer_reviewed","atlasIds":null},{"id":"c9ec2543-f689-4c53-a982-2aa7461fd670","title":"Corrupted but Correct: Why Vision-Language Models Lie to Themselves Internally","headline":null,"summary":"Researchers report that a targeted adversarial perturbation can drive a vision-language model's teacher-forced training loss for a fixed target caption near zero, while free generation on Qwen2.5-VL-7B-Instruct still yields the correct description. Using 200 held-out COCO images and a two-stage PGD attack, they localise the gap to one autoregressive step and across the 28 LLM decoder layers, finding that the language decoder, not the visual encoder, differentially arbitrates which images are corrupted (linear probe AUC=0.858, flagged for circularity). The authors argue that adversarial robustness in autoregressive VLMs depends mainly on the language decoder's prior.","sourceUrl":"https://arxiv.org/abs/2610.03445v1","publishedAt":"2026-10-02T15:27:28.000Z","severity":"info","cvssSeverity":null,"cvssScore":null,"labels":["security","research"],"issueType":"research","cveId":null,"cweIds":null,"affectedPackages":null,"affectedVendors":[],"affectedVendorsRaw":["Qwen2.5-VL-7B-Instruct"],"classifierModel":"claude-haiku-5-5","classifierPromptVersion":"v4","solution":"N/A -- no mitigation discussed in source.","attackType":["model_evasion"],"cvssVector":null,"attackVector":null,"attackComplexity":null,"privilegesRequired":null,"userInteraction":null,"exploitMaturity":null,"epssScore":null,"epssCheckedAt":null,"kevDateAdded":null,"advisoryAliases":null,"affectedPackagesSource":null,"patchAvailable":null,"disclosureDate":"2026-10-02T15:27:28.000Z","capecIds":null,"crossRefCount":0,"attackSophistication":"advanced","impactType":["integrity","safety"],"aiComponentTargeted":"model","llmSpecific":true,"classifierConfidence":0.93,"researchCategory":"preprint","atlasIds":null},{"id":"d48c9a06-0833-4393-8846-bf30a39b45eb","title":"It Takes Little to Rewrite Perception: Targeted Semantic Substitution in Vision-Language Models at $ε\\leq 4/255$","headline":null,"summary":"Researchers show that targeted semantic substitution can make vision-language models (VLMs) perceive a different target than the source image within an adversarial perturbation budget of ε ≤ 4/255, a range where prior representation-alignment attacks achieved limited success. Under a white-box threat model, the source image's token streams are aligned with the target's in the victim VLM's post-merger token space. Under a strict success criterion, complete replacement reaches 38% on images at ε = 4/255 and 35.9% on video at ε = 1/255.","sourceUrl":"https://arxiv.org/abs/2609.38298v2","publishedAt":"2026-09-29T17:52:39.000Z","severity":"info","cvssSeverity":null,"cvssScore":null,"labels":["security","research"],"issueType":"research","cveId":null,"cweIds":null,"affectedPackages":null,"affectedVendors":[],"affectedVendorsRaw":["Vision Language Models","VLMs"],"classifierModel":"claude-haiku-5-5","classifierPromptVersion":"v4","solution":"N/A -- no mitigation discussed in source.","attackType":["model_evasion"],"cvssVector":null,"attackVector":null,"attackComplexity":null,"privilegesRequired":null,"userInteraction":null,"exploitMaturity":null,"epssScore":null,"epssCheckedAt":null,"kevDateAdded":null,"advisoryAliases":null,"affectedPackagesSource":null,"patchAvailable":null,"disclosureDate":"2026-09-29T17:52:39.000Z","capecIds":null,"crossRefCount":0,"attackSophistication":"advanced","impactType":["integrity","safety"],"aiComponentTargeted":"model","llmSpecific":true,"classifierConfidence":0.93,"researchCategory":"preprint","atlasIds":null},{"id":"31f9cb1e-b02c-4e21-8309-397af3f8b450","title":"Selective Channel Restoration for Backdoored Vision-Language Models","headline":null,"summary":"Researchers propose Perturb-Select-Restore (PSR), a post-training defense against backdoors in vision-language models (VLMs) implanted through poisoned fine-tuning data. PSR performs sparse updates to the projection interface and adds no computation at inference time. The authors report that backdoored VLM projectors are more sensitive to bounded perturbations than clean ones, a property they call projection fragility, and PSR reduces attack success rates to near zero while preserving clean-task performance.","sourceUrl":"https://arxiv.org/abs/2609.37759v1","publishedAt":"2026-09-29T15:04:28.000Z","severity":"info","cvssSeverity":null,"cvssScore":null,"labels":["security","research"],"issueType":"research","cveId":null,"cweIds":null,"affectedPackages":null,"affectedVendors":[],"affectedVendorsRaw":["Vision-language models (VLMs)"],"classifierModel":"claude-haiku-5-5","classifierPromptVersion":"v4","solution":"PSR, a post-training defense that identifies the output channels most sensitive to perturbations in each projection layer of a backdoored VLM and restores their parameters to the corresponding pretrained values.","attackType":["other"],"cvssVector":null,"attackVector":null,"attackComplexity":null,"privilegesRequired":null,"userInteraction":null,"exploitMaturity":null,"epssScore":null,"epssCheckedAt":null,"kevDateAdded":null,"advisoryAliases":null,"affectedPackagesSource":null,"patchAvailable":null,"disclosureDate":"2026-09-29T15:04:28.000Z","capecIds":null,"crossRefCount":0,"attackSophistication":"advanced","impactType":["integrity"],"aiComponentTargeted":"model","llmSpecific":false,"classifierConfidence":0.95,"researchCategory":"preprint","atlasIds":null},{"id":"1b441a3e-1f57-47e0-b4e9-4d6d2063d8b4","title":"Empirical Analysis of Goal Hijacking in Large Vision-Language Models via Visual Prompt Injection","headline":null,"summary":"Researchers study visual prompt injection (VPI), where instructions embedded in input images are followed by large vision-language models (LVLMs). They propose \"goal hijacking via visual prompt injection\" (GHVPI), which redirects an LVLM from its original task to an attacker-specified one. Their quantitative analysis reports an attack success rate of 15.8% against GPT-4V, and they find GHVPI success depends on the character recognition and instruction-following capabilities of LVLMs.","sourceUrl":"https://doi.org/10.1007/s00354-026-00334-8","publishedAt":"2026-09-28T00:00:00.000Z","severity":"info","cvssSeverity":null,"cvssScore":null,"labels":["security","research"],"issueType":"research","cveId":null,"cweIds":null,"affectedPackages":null,"affectedVendors":["OpenAI"],"affectedVendorsRaw":["GPT-4V"],"classifierModel":"claude-haiku-5-5","classifierPromptVersion":"v4","solution":"N/A -- no mitigation discussed in source.","attackType":["prompt_injection"],"cvssVector":null,"attackVector":null,"attackComplexity":null,"privilegesRequired":null,"userInteraction":null,"exploitMaturity":null,"epssScore":null,"epssCheckedAt":null,"kevDateAdded":null,"advisoryAliases":null,"affectedPackagesSource":null,"patchAvailable":null,"disclosureDate":"2026-09-28T00:00:00.000Z","capecIds":null,"crossRefCount":0,"attackSophistication":"moderate","impactType":["integrity","safety"],"aiComponentTargeted":"model","llmSpecific":true,"classifierConfidence":0.95,"researchCategory":"peer_reviewed","atlasIds":null},{"id":"c68dce8c-2a1b-43e5-80aa-767472311f14","title":"ThreatsDay: Android Spyware, PLC Attacks, AI Image Prompt Injection + 12 More Stories","headline":null,"summary":"GitHub will begin rejecting command-line support bundle uploads from older GHES appliances starting August 18, 2026, unless they are patched. A separate npm package, @copilot-mcp/apex, acts as a postinstall dropper that installs a macOS infostealer, and a fake VS Code extension, \"Markdown All Pro\", impersonates Markdown All in One to beacon machine details and fetch remote payloads.","sourceUrl":"https://thehackernews.com/2026/07/threatsday-android-spyware-plc-attacks.html","publishedAt":"2026-07-23T15:02:07.000Z","severity":"medium","cvssSeverity":null,"cvssScore":null,"labels":["security","research"],"issueType":"news","cveId":null,"cweIds":null,"affectedPackages":null,"affectedVendors":[],"affectedVendorsRaw":[],"classifierModel":"claude-haiku-5-5","classifierPromptVersion":"v4","solution":"To avoid disruption when submitting support bundles, update your GHES instance to the latest patch release available for your current version line. At minimum, the required patch versions are: 3.21.3, 3.20.5, 3.19.9, 3.18.12, and 3.17.18.","attackType":["prompt_injection"],"cvssVector":null,"attackVector":null,"attackComplexity":null,"privilegesRequired":null,"userInteraction":null,"exploitMaturity":null,"epssScore":null,"epssCheckedAt":null,"kevDateAdded":null,"advisoryAliases":null,"affectedPackagesSource":null,"patchAvailable":null,"disclosureDate":"2026-07-23T15:02:07.000Z","capecIds":null,"crossRefCount":0,"attackSophistication":"moderate","impactType":["confidentiality","integrity"],"aiComponentTargeted":"agent","llmSpecific":true,"classifierConfidence":0.6,"researchCategory":null,"atlasIds":null},{"id":"0433778e-037d-49f3-8374-efe775a41073","title":"Enhancing Targeted Adversarial Attacks on Large Vision-Language Models via Intermediate Projector","headline":null,"summary":"Researchers show that black-box targeted attacks on Large Vision-Language Models can be made more precise by working through the projector, a semantic bridge between vision and language. They propose Intermediate Projector Guided Attack (IPGA), which aligns Q-Former query outputs with a target and transfers across models, and IPGA with Residual Query Alignment (IPGA-R), which also preserves non-target content for fine-grained edits. The authors report that IPGA beats baselines on global attacks, IPGA-R wins on fine-grained attacks, and the method transfers to Google Gemini and OpenAI GPT.","sourceUrl":"http://ieeexplore.ieee.org/document/11557371","publishedAt":"2026-06-10T13:17:33.000Z","severity":"info","cvssSeverity":null,"cvssScore":null,"labels":["security","research"],"issueType":"research","cveId":null,"cweIds":null,"affectedPackages":null,"affectedVendors":["Google","OpenAI"],"affectedVendorsRaw":["Google Gemini","OpenAI GPT","Vision-Language Models","Q-Former"],"classifierModel":"claude-haiku-5-5","classifierPromptVersion":"v4","solution":"N/A -- no mitigation discussed in source.","attackType":["model_evasion"],"cvssVector":null,"attackVector":null,"attackComplexity":null,"privilegesRequired":null,"userInteraction":null,"exploitMaturity":null,"epssScore":null,"epssCheckedAt":null,"kevDateAdded":null,"advisoryAliases":null,"affectedPackagesSource":null,"patchAvailable":null,"disclosureDate":"2026-06-10T13:17:33.000Z","capecIds":null,"crossRefCount":0,"attackSophistication":"advanced","impactType":["integrity","safety"],"aiComponentTargeted":"model","llmSpecific":true,"classifierConfidence":0.95,"researchCategory":"peer_reviewed","atlasIds":null},{"id":"61471733-33d2-42ec-9794-c4032cf882d7","title":"Introducing Gemma 4 12B: a unified, encoder-free multimodal model","headline":null,"summary":"Google DeepMind introduced Gemma 4 12B, a mid-sized model with native audio inputs designed to run locally on laptops with 16GB of VRAM or unified memory. Its encoder-free architecture feeds vision and audio inputs directly into the LLM backbone, and the model is released under an Apache 2.0 license with Multi-Token Prediction drafters to reduce latency.","sourceUrl":"https://deepmind.google/blog/introducing-gemma-4-12b-a-unified-encoder-free-multimodal-model/","publishedAt":"2026-06-09T14:10:19.000Z","severity":"info","cvssSeverity":null,"cvssScore":null,"labels":["industry","research"],"issueType":"news","cveId":null,"cweIds":null,"affectedPackages":null,"affectedVendors":["Google"],"affectedVendorsRaw":["Gemma 4 12B","Gemma 4","Google DeepMind","Ollama","LM Studio","Hugging Face","llama.cpp","MLX","SGLang","vLLM","Unsloth"],"classifierModel":"claude-haiku-5-5","classifierPromptVersion":"v4","solution":"N/A -- no mitigation discussed in source.","attackType":[],"cvssVector":null,"attackVector":null,"attackComplexity":null,"privilegesRequired":null,"userInteraction":null,"exploitMaturity":null,"epssScore":null,"epssCheckedAt":null,"kevDateAdded":null,"advisoryAliases":null,"affectedPackagesSource":null,"patchAvailable":null,"disclosureDate":"2026-06-09T14:10:19.000Z","capecIds":null,"crossRefCount":0,"attackSophistication":"moderate","impactType":null,"aiComponentTargeted":null,"llmSpecific":true,"classifierConfidence":0.9,"researchCategory":null,"atlasIds":null},{"id":"92a9bbc4-55cc-4d0c-bf11-8209090b32bb","title":"PVLM: Parsing-Aware Vision-Language Model With Dynamic Contrastive Learning for Zero-Shot Deepfake Attribution","headline":null,"summary":"PVLM is a parsing-aware vision-language model with dynamic contrastive learning for zero-shot deepfake attribution, meaning it traces forged faces to generators not seen in training, including diffusion models. The authors build a fine-grained ZS-DFA benchmark and use face parsing to exploit differences in how GAN and diffusion generators preserve source facial attributes. They report that the model exceeds the state of the art on the ZS-DFA benchmark across various protocol evaluations.","sourceUrl":"http://ieeexplore.ieee.org/document/11520180","publishedAt":"2026-05-14T13:17:22.000Z","severity":"info","cvssSeverity":null,"cvssScore":null,"labels":["security","research"],"issueType":"research","cveId":null,"cweIds":null,"affectedPackages":null,"affectedVendors":[],"affectedVendorsRaw":[],"classifierModel":"claude-haiku-5-5","classifierPromptVersion":"v4","solution":"N/A -- no mitigation discussed in source.","attackType":[],"cvssVector":null,"attackVector":null,"attackComplexity":null,"privilegesRequired":null,"userInteraction":null,"exploitMaturity":null,"epssScore":null,"epssCheckedAt":null,"kevDateAdded":null,"advisoryAliases":null,"affectedPackagesSource":null,"patchAvailable":null,"disclosureDate":"2026-05-14T13:17:22.000Z","capecIds":null,"crossRefCount":0,"attackSophistication":"advanced","impactType":null,"aiComponentTargeted":"model","llmSpecific":false,"classifierConfidence":0.85,"researchCategory":"peer_reviewed","atlasIds":null},{"id":"3b53ff38-0705-4615-b148-cda6b66f05dd","title":"Privacy-preserving for user-uploaded images and text in Vision-Language Models","headline":null,"summary":null,"sourceUrl":"https://www.sciencedirect.com/science/article/pii/S0167404826001070?dgcid=rss_sd_all","publishedAt":"2026-04-28T12:01:26.173Z","severity":"info","cvssSeverity":null,"cvssScore":null,"labels":["privacy","research"],"issueType":"research","cveId":null,"cweIds":null,"affectedPackages":null,"affectedVendors":[],"affectedVendorsRaw":["Vision-Language Models"],"classifierModel":"claude-haiku-5-5","classifierPromptVersion":"v4","solution":"N/A -- no mitigation discussed in source.","attackType":[],"cvssVector":null,"attackVector":null,"attackComplexity":null,"privilegesRequired":null,"userInteraction":null,"exploitMaturity":null,"epssScore":null,"epssCheckedAt":null,"kevDateAdded":null,"advisoryAliases":null,"affectedPackagesSource":null,"patchAvailable":null,"disclosureDate":null,"capecIds":null,"crossRefCount":0,"attackSophistication":"moderate","impactType":["confidentiality"],"aiComponentTargeted":"model","llmSpecific":true,"classifierConfidence":0.85,"researchCategory":"peer_reviewed","atlasIds":null},{"id":"2dc370a4-9951-48d1-b1c4-90788c133907","title":"VLBiasBench: A Comprehensive Benchmark for Evaluating Bias in Large Vision-Language Model","headline":null,"summary":"VLBiasBench is a benchmark for measuring social bias in Large Vision-Language Models (LVLMs). It covers nine bias categories plus two intersectional ones (race × gender and race × social economic status). The dataset has 128,342 samples built from 46,848 images generated with Stable Diffusion XL, and the authors evaluated 15 open-source models and two closed-source models.","sourceUrl":"http://ieeexplore.ieee.org/document/11481174","publishedAt":"2026-04-14T13:16:42.000Z","severity":"info","cvssSeverity":null,"cvssScore":null,"labels":["research","safety"],"issueType":"research","cveId":null,"cweIds":null,"affectedPackages":null,"affectedVendors":[],"affectedVendorsRaw":["Large Vision-Language Models (LVLMs)","Stable Diffusion XL"],"classifierModel":"claude-haiku-5-5","classifierPromptVersion":"v4","solution":"N/A -- no mitigation discussed in source.","attackType":[],"cvssVector":null,"attackVector":null,"attackComplexity":null,"privilegesRequired":null,"userInteraction":null,"exploitMaturity":null,"epssScore":null,"epssCheckedAt":null,"kevDateAdded":null,"advisoryAliases":null,"affectedPackagesSource":null,"patchAvailable":null,"disclosureDate":"2026-04-14T13:16:42.000Z","capecIds":null,"crossRefCount":0,"attackSophistication":"moderate","impactType":["safety"],"aiComponentTargeted":"model","llmSpecific":false,"classifierConfidence":0.9,"researchCategory":"peer_reviewed","atlasIds":null},{"id":"9f6bc978-82da-4763-bf3e-9ea1d107a480","title":"SMInject: Specious Malignant Injection Attacks With Semantically-Enhanced Tokens in Cross-Modal Retrieval","headline":null,"summary":"SMInject is a new injection attack framework against pre-trained multimodal models used in cross-modal retrieval. It generates deceptive injections that combine concepts through causal correlation across modalities, and aligns them in the encoded embedding space to boost effectiveness. On representative multimodal models, it reportedly achieves over 14% higher attack success rate and 6% higher Hit@5 than state-of-the-art methods while preserving overall model utility.","sourceUrl":"http://ieeexplore.ieee.org/document/11433760","publishedAt":"2026-03-13T13:17:13.000Z","severity":"info","cvssSeverity":null,"cvssScore":null,"labels":["security","research"],"issueType":"research","cveId":null,"cweIds":null,"affectedPackages":null,"affectedVendors":[],"affectedVendorsRaw":[],"classifierModel":"claude-haiku-5-5","classifierPromptVersion":"v4","solution":"N/A -- no mitigation discussed in source.","attackType":["other"],"cvssVector":null,"attackVector":null,"attackComplexity":null,"privilegesRequired":null,"userInteraction":null,"exploitMaturity":null,"epssScore":null,"epssCheckedAt":null,"kevDateAdded":null,"advisoryAliases":null,"affectedPackagesSource":null,"patchAvailable":null,"disclosureDate":"2026-03-13T13:17:13.000Z","capecIds":null,"crossRefCount":0,"attackSophistication":"advanced","impactType":["integrity"],"aiComponentTargeted":"model","llmSpecific":false,"classifierConfidence":0.95,"researchCategory":"peer_reviewed","atlasIds":null},{"id":"9f3d007d-db88-45d1-8748-07760b16ecfa","title":"Are Large Vision-Language Models Robust to Adversarial Visual Transformations?","headline":null,"summary":"Researchers assess how robust large vision-language models (LVLMs) are to adversarial visual transformations, a simpler attack style than optimizing perturbations or manipulating prompts. They test LVLM resilience across all possible transformation operations and find that combining the most harmful transformations yields more effective attacks. They also introduce adversarial learning of visual transformations, which applies malicious transformations to raw images via gradient approximation to improve attack effectiveness and imperceptibility.","sourceUrl":"http://ieeexplore.ieee.org/document/11421907","publishedAt":"2026-03-05T13:17:20.000Z","severity":"info","cvssSeverity":null,"cvssScore":null,"labels":["security","research"],"issueType":"research","cveId":null,"cweIds":null,"affectedPackages":null,"affectedVendors":[],"affectedVendorsRaw":["large vision-language models (LVLMs)"],"classifierModel":"claude-haiku-5-5","classifierPromptVersion":"v4","solution":"N/A -- no mitigation discussed in source.","attackType":["model_evasion"],"cvssVector":null,"attackVector":null,"attackComplexity":null,"privilegesRequired":null,"userInteraction":null,"exploitMaturity":null,"epssScore":null,"epssCheckedAt":null,"kevDateAdded":null,"advisoryAliases":null,"affectedPackagesSource":null,"patchAvailable":null,"disclosureDate":"2026-03-05T13:17:20.000Z","capecIds":null,"crossRefCount":0,"attackSophistication":"moderate","impactType":["integrity","safety"],"aiComponentTargeted":"model","llmSpecific":false,"classifierConfidence":0.95,"researchCategory":"peer_reviewed","atlasIds":null},{"id":"02c3456b-4ef6-4aa2-b53a-d515dd4219f6","title":"Evaluating and Mitigating Relationship Hallucinations in Large Vision-Language Models","headline":null,"summary":"R-Bench is a new benchmark for measuring hallucinations about relationships between objects in Large Vision-Language Models, using image-level questions on whether a relationship exists and instance-level questions on local visual comprehension. The authors trace these hallucinations to relationship-relationship, subject-relationship and relationship-object co-occurrences, worsened by long-tail distributions in visual datasets. They report that region-level image-text alignment reduces them and propose a baseline, Region-Aware Alignment Mitigation (RA$^{2}$2M), that directs model attention to relevant regions.","sourceUrl":"http://ieeexplore.ieee.org/document/11371674","publishedAt":"2026-02-03T13:17:37.000Z","severity":"info","cvssSeverity":null,"cvssScore":null,"labels":["research","safety"],"issueType":"research","cveId":null,"cweIds":null,"affectedPackages":null,"affectedVendors":[],"affectedVendorsRaw":[],"classifierModel":"claude-haiku-5-5","classifierPromptVersion":"v4","solution":"Region-level image-text alignment helps mitigate relationship hallucinations; the authors propose RA$^{2}$2M (Region-Aware Alignment Mitigation) as a new baseline that enhances model attention to relevant regions.","attackType":[],"cvssVector":null,"attackVector":null,"attackComplexity":null,"privilegesRequired":null,"userInteraction":null,"exploitMaturity":null,"epssScore":null,"epssCheckedAt":null,"kevDateAdded":null,"advisoryAliases":null,"affectedPackagesSource":null,"patchAvailable":null,"disclosureDate":"2026-02-03T13:17:37.000Z","capecIds":null,"crossRefCount":0,"attackSophistication":"moderate","impactType":["integrity"],"aiComponentTargeted":"model","llmSpecific":true,"classifierConfidence":0.9,"researchCategory":"peer_reviewed","atlasIds":null},{"id":"75bab064-0585-4de1-bb0f-c9d0bfa0e4ec","title":"NAP-Tuning: Neural Augmented Prompt Tuning for Adversarially Robust Vision-Language Models","headline":null,"summary":"NAP-Tuning extends Adversarial Prompt Tuning (AdvPT) for vision-language models such as CLIP by adding a multi-modal, multi-layer prompting framework. Its core is a Neural Augmentor that uses TokenRefiners, lightweight modules that reconstruct purified features through residual connections to correct adversarial distortions in feature space. Under the AutoAttack benchmark it outperforms the strongest baselines by 32.3% on ViT-B16 and 31.3% on ViT-B32 while keeping competitive clean accuracy.","sourceUrl":"http://ieeexplore.ieee.org/document/11368741","publishedAt":"2026-01-30T13:17:15.000Z","severity":"info","cvssSeverity":null,"cvssScore":null,"labels":["security","research"],"issueType":"research","cveId":null,"cweIds":null,"affectedPackages":null,"affectedVendors":[],"affectedVendorsRaw":["CLIP","ViT-B16","ViT-B32"],"classifierModel":"claude-haiku-5-5","classifierPromptVersion":"v4","solution":"N/A -- no mitigation discussed in source.","attackType":["model_evasion"],"cvssVector":null,"attackVector":null,"attackComplexity":null,"privilegesRequired":null,"userInteraction":null,"exploitMaturity":null,"epssScore":null,"epssCheckedAt":null,"kevDateAdded":null,"advisoryAliases":null,"affectedPackagesSource":null,"patchAvailable":null,"disclosureDate":"2026-01-30T13:17:15.000Z","capecIds":null,"crossRefCount":0,"attackSophistication":"advanced","impactType":["integrity"],"aiComponentTargeted":"model","llmSpecific":false,"classifierConfidence":0.95,"researchCategory":"peer_reviewed","atlasIds":null}],"meta":{"total":18,"limit":20,"offset":0}}