Adversarial machine learning
Attacks on how models learn and decide: adversarial examples, evasion, data poisoning and backdoors in trained models.
- All items
- 104
- Last 90 days
- 43
- Change
- +79%vs 24 before
Items per month
| Month | Items |
|---|---|
| May 2025 | 0 |
| Jun 2025 | 0 |
| Jul 2025 | 0 |
| Aug 2025 | 1 |
| Sep 2025 | 1 |
| Oct 2025 | 1 |
| Nov 2025 | 1 |
| Dec 2025 | 6 |
| Jan 2026 | 4 |
| Feb 2026 | 6 |
| Mar 2026 | 11 |
| Apr 2026 | 8 |
| May 2026 | 10 |
| Jun 2026 | 9 |
| Jul 2026 | 8 |
| Aug 2026 | 3 |
| Sep 2026 | 23 |
| Oct 2026 | 9 |
97 items
Detecting Adversarial Images through Response Profiles of Vision-Language Models
Oct 7, 2026InfoResearchPreprintSecurityResearchThe paper proposes a detector that identifies adversarial images for frozen vision-language models by profiling how an image responds to a set of general semantic prompts. The profile combines category-level statistics, prompt relationships, deviations from clean reference distributions, and stability under weak image transformations, and a lightweight classifier labels each input while the VLM stays fixed. Evaluated across multiple datasets, CLIP-style backbones and several attack families, the detector discriminates strongly in attack-specific settings and retains substantial performance on unseen attacks.
Arxiv (cs.RO + cs.CV security)GraphRectify: Graph-Based Transfer of Adversarial Example Detectors Across Neural Networks
Oct 7, 2026InfoResearchPreprintSecurityResearchGraphRectify is a graph-based framework that transfers adversarial image detectors from one classifier backbone to another. It learns a structured representation of intermediate classifier features and adapts features from a new backbone to the detector trained on the original model. Across the evaluation matrix, it achieves higher aggregate ROC-AUC than training a detector from scratch on the new backbone, with the largest gains between different backbone families when sufficient data are available.
Arxiv (cs.RO + cs.CV security)Transferable Spatial Temporal Coherence Adversarial Attack on Black-Box Vision Language Models for Autonomous Driving
Oct 6, 2026InfoResearchPreprintSecurityResearchResearchers introduce STCA (Spatial Temporal Coherence Adversarial Attack), a black-box method that perturbs driving video to fool Vision Language Models. The attack selects semantically important frames with caption guidance, applies a spatial perturbation that preserves high SSIM, then disrupts cross-frame temporal coherence with a motion-guided mask. Tested on BDD100K and nuScenes against Video LLaVA-7B, Qwen2.5-VL-7B and Dolphin, the spatial attack reaches a high ASR while keeping SSIM high, showing these models remain highly susceptible.
Arxiv (cs.RO + cs.CV security)TAPDreamer: Transferable Adversarial Patches for World Action Models
Oct 5, 2026LowResearchPreprintSecurityResearchTAPDreamer is an attack on world action models that builds a fixed local adversarial patch using only a public encoder, with no queries to the target policy. The patch, covering about 6.5% of the input, transfers across tasks and action architectures. In closed-loop tests it cut FastWAM's success rate from 97.7% to 0.0% on 40 LIBERO tasks and from 90.86% to 0.0% on 50 RoboTwin tasks, and it also lowered success on two DreamWAM configurations and on Motus.
Arxiv (cs.RO + cs.CV security)A comprehensive study of cross-domain adversarial robustness and attack transferability in image-based malware detection and classification
Oct 4, 2026InfoResearchPeer-reviewedResearchSecurityThis paper presents a framework for evaluating adversarial robustness and attack transferability in image-based deep learning models for malware detection and classification. The authors apply image-domain attacks from FGSM to AutoAttack and test whether binary-domain manipulations remain effective after conversion to an image representation. They report average attack success rates of 64.6% for FGSM and 98.8% for AutoAttack, and accuracy drops of up to 42% from transferred binary-domain manipulations.
OpenAlex (peer-reviewed AI security)Visual-Invariance-Augmented Feature Optimal Alignment for Transferable Adversarial Attacks against Closed-Source MLLMs
Oct 3, 2026InfoResearchPreprintSecurityResearchResearchers study how to make adversarial images transfer from open-source surrogate multimodal LLMs to closed-source MLLMs in black-box settings. They propose IAU-FOA, which aligns adversarial and target images at both global and patch-cluster levels using confidence-adaptive unbalanced optimal transport, plus visual-invariance augmentation that simulates exposure, contrast, illumination and color-temperature changes. The authors report that it consistently outperforms state-of-the-art transferable attack methods across open-source and closed-source MLLMs.
Arxiv (cs.RO + cs.CV security)A SHAP-guided heterogeneous ensemble defense framework for financial risk assessment under white-box adversarial attacks
Oct 3, 2026InfoResearchPeer-reviewedSecurityResearchResearchers evaluated a three-layer ensemble defense for deep learning credit risk models, combining MLP, ResNet-1D and TabTransformer architectures with PGD adversarial training and SHAP-based routing. On German Credit and Lending Club under FGSM, PGD and CW attacks across five seeds, the defended AUCs reached 0.758 and 0.723, and the default-class attack success rate fell from 0.52 to 0.19 and from 0.55 to 0.21. SHAP attribution consistency also improved, with Spearman correlation rising from 0.42 to 0.87 between clean and defended explanations.
Fix: The defense framework itself is the proposed mitigation: a three-layer heterogeneous ensemble with PGD adversarial training and SHAP-based routing. The source does not describe a separate patch, fixed version or configuration change.
OpenAlex (peer-reviewed AI security)Detect and Suppress: A Mechanistic Defense against Adversarial Patches in VLA Models
Oct 2, 2026InfoResearchPreprintSecurityResearchResearchers analyze Vision-Language-Action (VLA) models with a sparse autoencoder (SAE) and find an internal feature whose activation strongly correlates with adversarial patches. They suppress this feature at inference time only when a linear probe detects an attack, which improves robustness without fine-tuning the VLA. On LIBERO-10, conditional intervention raises success rate under intermittent attacks, while continuous intervention substantially degrades policy performance.
Fix: Suppress the identified SAE feature at inference time, applying the intervention only when a linear probe detects an attack. Avoid continuous application, which substantially degrades policy performance.
Arxiv (cs.RO + cs.CV security)Corrupted but Correct: Why Vision-Language Models Lie to Themselves Internally
Oct 2, 2026InfoResearchPreprintSecurityResearchResearchers report that a targeted adversarial perturbation can drive a vision-language model's teacher-forced training loss for a fixed target caption near zero, while free generation on Qwen2.5-VL-7B-Instruct still yields the correct description. Using 200 held-out COCO images and a two-stage PGD attack, they localise the gap to one autoregressive step and across the 28 LLM decoder layers, finding that the language decoder, not the visual encoder, differentially arbitrates which images are corrupted (linear probe AUC=0.858, flagged for circularity). The authors argue that adversarial robustness in autoregressive VLMs depends mainly on the language decoder's prior.
Arxiv (cs.RO + cs.CV security)Hard-label black-box model extraction attacks against network intrusion detection systems via generative adversarial networks
Sep 30, 2026InfoResearchPeer-reviewedSecurityResearchThe paper asks whether a network intrusion detection system's model can be extracted using only hard-label black-box queries. The source text provided gives only the title, publication date (December 2026), journal (Journal of Information Security and Applications, Volume 103), and authors (Donguk Min, Seungsoo Nam, Daeseon Choi), with no method description or findings.
Elsevier Security JournalsLet the Carrier Carry the Attack: Preserving the Subject in Adversarial Image Generation
Sep 30, 2026InfoResearchPreprintSecurityResearchResearchers propose a "carrier", a secondary visual element that gives unrestricted adversarial attacks an auxiliary region to work in, so the primary object (the subject) is distorted less. Under global classifier guidance, the carrier absorbs a larger share of normalized attack updates and improves cross-model transferability. Targeted attacks keep the personalized subject as the main content perceived by humans while misleading the classifier.
Arxiv (cs.RO + cs.CV security)Universal Cross-Prompt Adversarial Attacks on Promptable Concept Segmentation
Sep 30, 2026InfoResearchPreprintSecurityResearchResearchers present AdvPCS, a universal cross-prompt adversarial attack on Promptable Concept Segmentation in SAM3, which extends SAM-series models to concept-level prediction. The method combines min-max prompt optimization with perception deception and temporal memory misalignment attacks. A single universal adversarial perturbation (UAP) generalizes across frames from different videos and reduces the average mIoU of PCS models on the SA-CO dataset to below 5% under text prompts.
Arxiv (cs.RO + cs.CV security)Exploiting Vulnerabilities: Universal Adversarial Attacks on Vision-Language-Action Models in Robotics
Sep 30, 2026InfoResearchPreprintSecurityResearchResearchers propose a Universal Adversarial Object, a sphere with an optimized surface texture, that degrades the task success of Vision-Language-Action (VLA) models when placed in a robot's field of view. Their multi-level attack framework disrupts trajectory planning, task execution and action control, and was validated in simulated and real-world robotic settings. The object reduces average task success rates by 31.2% to 39.9% for two VLA models, Pi0 and RDT, with success rates dropping to near zero in complex scenarios.
Arxiv (cs.RO + cs.CV security)BadAction: Backdoor Attacks on Interactive Video Generation via Action-Guided Triggers
Sep 30, 2026InfoResearchPreprintSecurityResearchResearchers present the first systematic study of backdoor attacks against the interactivity of interactive video generation (IVG) models. Their method, BadAction, implants predefined motion patterns into action sequences so that triggered models produce frozen future frames that ignore later user actions, while benign inputs behave normally. The attack reaches average success rates of 91.0% with action-only triggers and 80.4% with multimodal triggers, and it bypasses existing backdoor detection methods.
Arxiv (cs.RO + cs.CV security)Efficient model stealing in data-free scenarios: An attack method via elite sample distillation
Sep 29, 2026InfoResearchPeer-reviewedSecurityResearchLiu, Wen and Yu published an article in the Journal of Information Security and Applications, Volume 103, dated December 2026. The source text provided contains only the publication metadata and author list, with no abstract, method or findings.
Elsevier Security JournalsSwitchable backdoor attack in vision transformers via progressive dual-token injection from shallow to deep layers
Sep 28, 2026InfoResearchPeer-reviewedSecurityResearchOpenAlex (peer-reviewed AI security)Exploring backdoor attack and defense algorithms in LLMS: Enhancing in-context learning security
Sep 28, 2026InfoResearchPeer-reviewedSecurityResearchThis paper shows that an attacker can manipulate LLM behavior by poisoning the demonstration context used in in-context learning, without fine-tuning the model. The authors present ICLAttack, a backdoor method that poisons demonstration examples or demonstration prompts, reporting a 95.0% average attack success rate on OPT models across three datasets. They also propose ICLDefense, which uses a lightweight auxiliary model and an ensemble-based strategy to refine LLM outputs and reduce attack success.
Fix: ICLDefense: a defense algorithm that utilizes model ensembles, employing a lightweight auxiliary model to refine LLM outputs through an ensemble-based strategy, which the source says substantially reduces the attack success rate compared to existing methods while preserving model performance.
OpenAlex (peer-reviewed AI security)Detection of Adversarial Attacks on Super-Resolvers Using Spectral Features
Sep 28, 2026InfoResearchPreprintSecurityResearchResearchers propose a spectral detection method for finding adversarial attacks embedded in super-resolution model weights, a preprocessing component of imaging pipelines. The method trains an XGBoost detector on the radially-averaged power spectral density and benchmarks it against magnitude- and phase-based Fourier spectrum detectors across training and cross-architecture scenarios. The proposed detector outperforms the comparison detectors in most scenarios, and high-frequency features prove most informative for detecting AdvSR attacks.
Arxiv (cs.RO + cs.CV security)IndirectAD: Practical Data Poisoning Attacks Against Recommender Systems for Item Promotion
Sep 28, 2026LowResearchPeer-reviewedSecurityResearchResearchers introduce IndirectAD, a data poisoning attack against recommender systems inspired by Trojan attacks on machine learning. The attack first promotes a trigger item, then transfers that advantage to a target item by creating co-occurrence data between them, which reduces the number of controlled accounts needed. Experiments on multiple datasets and recommender systems show noticeable impact with only 0.05% of a platform's user base.
IEEE Xplore (Security & AI Journals)Stealthy Physical Adversarial Attacks on Speaker Recognition via Near-Ultrasonic Perturbations
Sep 28, 2026LowResearchPeer-reviewedSecurityResearchAdvNup is a physical adversarial attack that spoofs deep neural network speaker recognition systems using near-ultrasonic perturbations played through commercial off-the-shelf speakers, avoiding the specialized hardware that earlier ultrasound attacks required. The authors use single-sideband modulation with low-pass filtering, a nonlinear frequency response model, and time-frequency masking to keep the adversarial signal intact through physical transmission. In simulated and physical experiments, it reached a 100% attack success rate for closed-set identification and over 90% for open-set identification in targeted attacks.
IEEE Xplore (Security & AI Journals)
Topic added 2026-10-09. An item belongs to this topic when its title matches one of the topic's patterns or its summary mentions the topic at least twice. Report a wrong match with the feedback button on the item.