Adversarial machine learning
Attacks on how models learn and decide: adversarial examples, evasion, data poisoning and backdoors in trained models.
- All items
- 108
- Last 90 days
- 47
- Change
- +114%vs 22 before
Items per month
| Month | Items |
|---|---|
| May 2025 | 0 |
| Jun 2025 | 0 |
| Jul 2025 | 0 |
| Aug 2025 | 1 |
| Sep 2025 | 1 |
| Oct 2025 | 1 |
| Nov 2025 | 1 |
| Dec 2025 | 6 |
| Jan 2026 | 4 |
| Feb 2026 | 6 |
| Mar 2026 | 11 |
| Apr 2026 | 8 |
| May 2026 | 10 |
| Jun 2026 | 9 |
| Jul 2026 | 8 |
| Aug 2026 | 3 |
| Sep 2026 | 24 |
| Oct 2026 | 12 |
108 items
Strong evasive backdoor attacks and an ensemble defense
Oct 7, 2026InfoResearchPeer-reviewedSecurityResearchResearchers present an ensemble defense that combines several backdoor detectors for deep neural network classifiers, so detection does not depend on a single backdoor mechanism. The ensemble also performs backdoor inversion, which indicates the nature of a detected attack. The paper also employs mixed clean/dirty-label backdoor poisoning, an X-to-X attack the authors describe as more surgical, evasive, and harder to detect than traditional dirty-label attacks.
OpenAlex (peer-reviewed AI security)GNN-Guided Selection of Benign-like Anomalies for Backdoor Attacks Against Network Intrusion Detection Systems
Oct 7, 2026LowResearchPeer-reviewedSecurityResearchThis paper studies a backdoor attack on AI-based network intrusion detection systems (NIDSs). The method uses the embedding space of a Graph Attention Network (GAT) to pick anomalous samples close to benign traffic, adds a benign-distribution trigger, and relabels them as benign before retraining TabNet, ACID, and AlertNet. Tests on NSL-KDD, CICIDS2017, and UNSW-NB15 show high attack success rates, but the gain over random selection is uneven and marginal or absent where the random baseline is already saturated.
OpenAlex (peer-reviewed AI security)Detecting Adversarial Images through Response Profiles of Vision-Language Models
Oct 7, 2026InfoResearchPreprintSecurityResearchThe paper proposes a detector that identifies adversarial images for frozen vision-language models by profiling how an image responds to a set of general semantic prompts. The profile combines category-level statistics, prompt relationships, deviations from clean reference distributions, and stability under weak image transformations, and a lightweight classifier labels each input while the VLM stays fixed. Evaluated across multiple datasets, CLIP-style backbones and several attack families, the detector discriminates strongly in attack-specific settings and retains substantial performance on unseen attacks.
Arxiv (cs.RO + cs.CV security)GraphRectify: Graph-Based Transfer of Adversarial Example Detectors Across Neural Networks
Oct 7, 2026InfoResearchPreprintSecurityResearchGraphRectify is a graph-based framework that transfers adversarial image detectors from one classifier backbone to another. It learns a structured representation of intermediate classifier features and adapts features from a new backbone to the detector trained on the original model. Across the evaluation matrix, it achieves higher aggregate ROC-AUC than training a detector from scratch on the new backbone, with the largest gains between different backbone families when sufficient data are available.
Arxiv (cs.RO + cs.CV security)Transferable Spatial Temporal Coherence Adversarial Attack on Black-Box Vision Language Models for Autonomous Driving
Oct 6, 2026InfoResearchPreprintSecurityResearchResearchers introduce STCA (Spatial Temporal Coherence Adversarial Attack), a black-box method that perturbs driving video to fool Vision Language Models. The attack selects semantically important frames with caption guidance, applies a spatial perturbation that preserves high SSIM, then disrupts cross-frame temporal coherence with a motion-guided mask. Tested on BDD100K and nuScenes against Video LLaVA-7B, Qwen2.5-VL-7B and Dolphin, the spatial attack reaches a high ASR while keeping SSIM high, showing these models remain highly susceptible.
Arxiv (cs.RO + cs.CV security)TAPDreamer: Transferable Adversarial Patches for World Action Models
Oct 5, 2026LowResearchPreprintSecurityResearchTAPDreamer is an attack on world action models that builds a fixed local adversarial patch using only a public encoder, with no queries to the target policy. The patch, covering about 6.5% of the input, transfers across tasks and action architectures. In closed-loop tests it cut FastWAM's success rate from 97.7% to 0.0% on 40 LIBERO tasks and from 90.86% to 0.0% on 50 RoboTwin tasks, and it also lowered success on two DreamWAM configurations and on Motus.
Arxiv (cs.RO + cs.CV security)A comprehensive study of cross-domain adversarial robustness and attack transferability in image-based malware detection and classification
Oct 4, 2026InfoResearchPeer-reviewedResearchSecurityThis paper presents a framework for evaluating adversarial robustness and attack transferability in image-based deep learning models for malware detection and classification. The authors apply image-domain attacks from FGSM to AutoAttack and test whether binary-domain manipulations remain effective after conversion to an image representation. They report average attack success rates of 64.6% for FGSM and 98.8% for AutoAttack, and accuracy drops of up to 42% from transferred binary-domain manipulations.
OpenAlex (peer-reviewed AI security)Visual-Invariance-Augmented Feature Optimal Alignment for Transferable Adversarial Attacks against Closed-Source MLLMs
Oct 3, 2026InfoResearchPreprintSecurityResearchResearchers study how to make adversarial images transfer from open-source surrogate multimodal LLMs to closed-source MLLMs in black-box settings. They propose IAU-FOA, which aligns adversarial and target images at both global and patch-cluster levels using confidence-adaptive unbalanced optimal transport, plus visual-invariance augmentation that simulates exposure, contrast, illumination and color-temperature changes. The authors report that it consistently outperforms state-of-the-art transferable attack methods across open-source and closed-source MLLMs.
Arxiv (cs.RO + cs.CV security)A SHAP-guided heterogeneous ensemble defense framework for financial risk assessment under white-box adversarial attacks
Oct 3, 2026InfoResearchPeer-reviewedSecurityResearchResearchers evaluated a three-layer ensemble defense for deep learning credit risk models, combining MLP, ResNet-1D and TabTransformer architectures with PGD adversarial training and SHAP-based routing. On German Credit and Lending Club under FGSM, PGD and CW attacks across five seeds, the defended AUCs reached 0.758 and 0.723, and the default-class attack success rate fell from 0.52 to 0.19 and from 0.55 to 0.21. SHAP attribution consistency also improved, with Spearman correlation rising from 0.42 to 0.87 between clean and defended explanations.
Fix: The defense framework itself is the proposed mitigation: a three-layer heterogeneous ensemble with PGD adversarial training and SHAP-based routing. The source does not describe a separate patch, fixed version or configuration change.
OpenAlex (peer-reviewed AI security)Detect and Suppress: A Mechanistic Defense against Adversarial Patches in VLA Models
Oct 2, 2026InfoResearchPreprintSecurityResearchResearchers analyze Vision-Language-Action (VLA) models with a sparse autoencoder (SAE) and find an internal feature whose activation strongly correlates with adversarial patches. They suppress this feature at inference time only when a linear probe detects an attack, which improves robustness without fine-tuning the VLA. On LIBERO-10, conditional intervention raises success rate under intermittent attacks, while continuous intervention substantially degrades policy performance.
Fix: Suppress the identified SAE feature at inference time, applying the intervention only when a linear probe detects an attack. Avoid continuous application, which substantially degrades policy performance.
Arxiv (cs.RO + cs.CV security)Corrupted but Correct: Why Vision-Language Models Lie to Themselves Internally
Oct 2, 2026InfoResearchPreprintSecurityResearchResearchers report that a targeted adversarial perturbation can drive a vision-language model's teacher-forced training loss for a fixed target caption near zero, while free generation on Qwen2.5-VL-7B-Instruct still yields the correct description. Using 200 held-out COCO images and a two-stage PGD attack, they localise the gap to one autoregressive step and across the 28 LLM decoder layers, finding that the language decoder, not the visual encoder, differentially arbitrates which images are corrupted (linear probe AUC=0.858, flagged for circularity). The authors argue that adversarial robustness in autoregressive VLMs depends mainly on the language decoder's prior.
Arxiv (cs.RO + cs.CV security)Fix Your Downsampling ASAP! Aliasing and Sinc Artifact Free Pooling in the Fourier Domain
Sep 30, 2026InfoResearchPeer-reviewedSecurityResearchResearchers show that CNN downsampling layers violate the Sampling Theorem, and that this aliasing correlates with vulnerability to adversarial attacks and distribution shifts. They propose Frequency Low Cut Pooling (FLC Pooling) and its extension Aliasing and Sinc Artifact-free Pooling (ASAP), an alias-free downsampling operation in the frequency domain that also removes sinc-interpolation artifacts. On ImageNet-1k, ImageNet-C and CIFAR across several CNN architectures, networks using these methods showed greater robustness to common corruptions and adversarial attacks while keeping clean accuracy close to baseline models.
OpenAlex (peer-reviewed AI security)Hard-label black-box model extraction attacks against network intrusion detection systems via generative adversarial networks
Sep 30, 2026InfoResearchPeer-reviewedSecurityResearchThe paper asks whether a network intrusion detection system's model can be extracted using only hard-label black-box queries. The source text provided gives only the title, publication date (December 2026), journal (Journal of Information Security and Applications, Volume 103), and authors (Donguk Min, Seungsoo Nam, Daeseon Choi), with no method description or findings.
Elsevier Security JournalsLet the Carrier Carry the Attack: Preserving the Subject in Adversarial Image Generation
Sep 30, 2026InfoResearchPreprintSecurityResearchResearchers propose a "carrier", a secondary visual element that gives unrestricted adversarial attacks an auxiliary region to work in, so the primary object (the subject) is distorted less. Under global classifier guidance, the carrier absorbs a larger share of normalized attack updates and improves cross-model transferability. Targeted attacks keep the personalized subject as the main content perceived by humans while misleading the classifier.
Arxiv (cs.RO + cs.CV security)Universal Cross-Prompt Adversarial Attacks on Promptable Concept Segmentation
Sep 30, 2026InfoResearchPreprintSecurityResearchResearchers present AdvPCS, a universal cross-prompt adversarial attack on Promptable Concept Segmentation in SAM3, which extends SAM-series models to concept-level prediction. The method combines min-max prompt optimization with perception deception and temporal memory misalignment attacks. A single universal adversarial perturbation (UAP) generalizes across frames from different videos and reduces the average mIoU of PCS models on the SA-CO dataset to below 5% under text prompts.
Arxiv (cs.RO + cs.CV security)Exploiting Vulnerabilities: Universal Adversarial Attacks on Vision-Language-Action Models in Robotics
Sep 30, 2026InfoResearchPreprintSecurityResearchResearchers propose a Universal Adversarial Object, a sphere with an optimized surface texture, that degrades the task success of Vision-Language-Action (VLA) models when placed in a robot's field of view. Their multi-level attack framework disrupts trajectory planning, task execution and action control, and was validated in simulated and real-world robotic settings. The object reduces average task success rates by 31.2% to 39.9% for two VLA models, Pi0 and RDT, with success rates dropping to near zero in complex scenarios.
Arxiv (cs.RO + cs.CV security)BadAction: Backdoor Attacks on Interactive Video Generation via Action-Guided Triggers
Sep 30, 2026InfoResearchPreprintSecurityResearchResearchers present the first systematic study of backdoor attacks against the interactivity of interactive video generation (IVG) models. Their method, BadAction, implants predefined motion patterns into action sequences so that triggered models produce frozen future frames that ignore later user actions, while benign inputs behave normally. The attack reaches average success rates of 91.0% with action-only triggers and 80.4% with multimodal triggers, and it bypasses existing backdoor detection methods.
Arxiv (cs.RO + cs.CV security)Efficient model stealing in data-free scenarios: An attack method via elite sample distillation
Sep 29, 2026InfoResearchPeer-reviewedSecurityResearchLiu, Wen and Yu published an article in the Journal of Information Security and Applications, Volume 103, dated December 2026. The source text provided contains only the publication metadata and author list, with no abstract, method or findings.
Elsevier Security JournalsVLM4Cluster: Benchmarking Deep Clustering In the Era of Vision-Language Pre-training
Sep 28, 2026InfoResearchPreprintResearchSecurityVLM4Cluster is a benchmark for image clustering with pre-trained vision-language models. It implements 17 methods across classical, deep, and language-assisted clustering and evaluates them on 20 datasets, including tests of adversarial robustness, distribution-shift generalization, and computational efficiency. The authors find that language-assisted clustering (LaIC) generally improves clustering performance and generalization, but its gains are less consistent on large-scale and fine-grained datasets, and language assistance does not systematically reduce sensitivity to adversarial perturbations.
Arxiv (cs.RO + cs.CV security)Switchable backdoor attack in vision transformers via progressive dual-token injection from shallow to deep layers
Sep 28, 2026InfoResearchPeer-reviewedSecurityResearchOpenAlex (peer-reviewed AI security)
Topic added 2026-10-09. An item belongs to this topic when its title matches one of the topic's patterns or its summary mentions the topic at least twice. Report a wrong match with the feedback button on the item.