Adversarial machine learning
Attacks on how models learn and decide: adversarial examples, evasion, data poisoning and backdoors in trained models.
- All items
- 108
- Last 90 days
- 47
- Change
- +114%vs 22 before
Items per month
| Month | Items |
|---|---|
| May 2025 | 0 |
| Jun 2025 | 0 |
| Jul 2025 | 0 |
| Aug 2025 | 1 |
| Sep 2025 | 1 |
| Oct 2025 | 1 |
| Nov 2025 | 1 |
| Dec 2025 | 6 |
| Jan 2026 | 4 |
| Feb 2026 | 6 |
| Mar 2026 | 11 |
| Apr 2026 | 8 |
| May 2026 | 10 |
| Jun 2026 | 9 |
| Jul 2026 | 8 |
| Aug 2026 | 3 |
| Sep 2026 | 24 |
| Oct 2026 | 12 |
108 items
Fraud-RLA: A Reinforcement Learning Adversarial Attack Against Credit Card Fraud Detection
Mar 13, 2026LowResearchPeer-reviewedSecurityResearchFraud-RLA is a new adversarial attack against credit card fraud detection systems that uses reinforcement learning to evade detection. It aims to maximize the amount stolen while requiring substantially less prior knowledge than competing methods. The authors report that it remains effective on a realistic fraud detection system under the constraints of their threat model.
IEEE Xplore (Security & AI Journals)Robustness Over Time: Understanding Adversarial Examples’ Effectiveness on Longitudinal Versions of Large Language Models
Mar 9, 2026InfoResearchPeer-reviewedResearchSafetyA longitudinal study tested the adversarial robustness of GPT, Llama and Qwen model families across successive versions, covering misclassification, jailbreak and hallucination. The authors found that LLM updates do not consistently improve robustness: a later GPT-3.5 version got worse on misclassification and hallucination despite better jailbreak resilience. GPT-4 and GPT-4o showed incrementally higher overall robustness, while larger Llama and Qwen models did not uniformly improve, and larger size did not reliably help.
IEEE Xplore (Security & AI Journals)DUAP: Disentanglement-Based Universal Adversarial Perturbations for Robust Multilingual Speech Privacy Protection
Mar 6, 2026InfoResearchPeer-reviewedSecurityPrivacyResearchDUAP is a privacy-preserving method that generates universal adversarial perturbations to counter OpenAI's Whisper multilingual ASR model, which can transcribe sensitive speech across languages. It uses a two-stage language attack: a Language Feature Disentanglement model produces adversarial examples, then gradient-based optimization disrupts Whisper's language identification module. Across three Whisper model sizes, DUAP reports WERs above 95% for English, 85% for other languages and 87% in physical settings, with SNRs from 40 dB down to above 17 dB.
IEEE Xplore (Security & AI Journals)Complementary Text-Guided Attention for Zero-Shot Adversarial Robustness
Mar 2, 2026InfoResearchPeer-reviewedSecurityResearchThe authors examine how adversarial perturbations shift text-guided attention in CLIP, a pre-trained vision-language model. They propose TGA-ZSR, which uses a Local Attention Refinement Module and a Global Attention Constraint Module, and Comp-TGA, which adds complementary class-prompt and non-class-prompt foreground attention. Across 16 datasets, TGA-ZSR and Comp-TGA improve zero-shot robust accuracy over the prior state of the art by 9.58% and 11.95%, respectively.
IEEE Xplore (Security & AI Journals)AdvDiffusion: Adversarial Patches Generation for Face Recognition With High Transferability in Physical Domain
Mar 2, 2026InfoResearchPeer-reviewedSecurityResearchThe paper presents AdvDiffusion, a method for generating adversarial face patches that can be pasted onto a face to cause face recognition models to make false identity judgments. It first selects a patch region using facial gradient maps, then adds noise to an image and denoises it with a pre-trained diffusion model, guided by an adversarial loss. Experiments report effective attacks on black-box face recognition models in both digital and physical settings, with better robustness to complex physical environments than some state-of-the-art methods.
IEEE Xplore (Security & AI Journals)PPOM-Attack: A Substitute Model-Free Perturbation Prediction and Optimization Method for Black-Box Adversarial Attack Against Face Recognition
Feb 23, 2026LowResearchPeer-reviewedSecurityResearchPPOM-Attack is a black-box attack against face recognition that avoids substitute models by taking feedback directly from the target model. It uses a proximal policy optimization (PPO)-based agent to predict and disturb perturbation regions in a face image, and a minimum brightness offsets method to keep adversarial images high quality. The authors report a 21.7% average gain in attack success rate over state-of-the-art FR attacks across seven FR models.
IEEE Xplore (Security & AI Journals)The democratization of AI data poisoning and how to protect your organization
Feb 13, 2026InfoNewsSecuritySafetyResearch cited in the article indicates that roughly 250 documents or images can distort the behavior of a large language model, far below earlier assumptions of thousands or millions of corrupted data points. Online communities have already begun seeding fabricated facts to influence LLM training data, and a Purdue, Texas A&M and UT Austin team found that junk data causes capability decay that clean data added later did not fully reverse.
Fix: The source recommends establishing a clean, validated "gold" version of the trusted model before deployment, to serve as a baseline for anomaly checks and as a restore point if outputs become unexpected or drift appears. It also calls for security controls to detect poisoning attacks, but does not specify them.
CSO OnlinePractical and Flexible Backdoor Attack Against Deep Learning Models via Shell Code Injection
Feb 9, 2026InfoResearchPeer-reviewedSecurityResearchResearchers propose shell code injection (SCI), a non-poisoning backdoor attack framework against deep learning models that needs no training. It uses a logic-driven backdoor shell with trigger consistency verification and short-circuit code packaging to force wrong predictions, and an LLM-assisted module generates attack target code adaptively. Experiments report almost 100% attack success rate (ASR) across various settings.
IEEE Xplore (Security & AI Journals)AdvScan: Black-Box Adversarial Example Detection at Runtime Through Power Analysis
Feb 9, 2026InfoResearchPeer-reviewedSecurityResearchAdvScan is a black-box method for detecting adversarial examples on TinyML models at runtime by analyzing power consumption. It builds a baseline of power signatures from benign inputs and applies a one-sample t-test to flag inputs whose signatures deviate significantly. Evaluated on three MLPerf Tiny models across two STM32 microcontrollers and 318,400 test inputs, it detected 99.984% of adversarial examples, with 40 false negatives and no false positives.
IEEE Xplore (Security & AI Journals)PROTheft: A Projector-Based Model Extraction Attack in the Physical World
Feb 6, 2026InfoResearchPeer-reviewedSecurityResearchPROTheft is a model extraction attack that extends digital-domain attacks to vision-based devices in the physical world. An attacker uses a projector placed in front of the on-board camera to feed attack samples to a black-box target, and a simulation module compensates for detail loss in the digital-to-physical-to-digital transformation. On a public autonomous driving dataset, the authors report over 80% fidelity with the target model and an mAP 50 above 0.85.
IEEE Xplore (Security & AI Journals)Allies Teach Better Than Enemies: Inverse Adversaries for Robust Knowledge Distillation
Feb 3, 2026InfoResearchPeer-reviewedResearchSecurityThis paper proposes an adversarially robust knowledge distillation method that compresses a large robust teacher model into a lightweight student. It refines inputs into inverse adversarial examples by reversing the sign of the adversarial perturbation, then applies gradient matching between teacher and student on those inputs, plus a weight-space disruption strategy. On ImageNet, the authors report roughly 3.8% gains over prior methods in both clean and robust accuracy.
IEEE Xplore (Security & AI Journals)A Wolf in Sheep’s Clothing: Unveiling a Stealthy Backdoor Attack in Subgraph Federated Learning
Jan 28, 2026InfoResearchPeer-reviewedSecurityResearchResearchers propose BEEF, an end-to-end backdoor attack against subgraph federated learning for node classification, where subgraphs are distributed across devices. BEEF uses a trigger generator trained jointly with the backdoored model, crafting adversarial perturbations as triggers that cause misclassification without changing model parameters. Evaluations across eight datasets, four models, five attacks and six aggregation methods report effectiveness against GNNs with minimal impact on normal data performance.
IEEE Xplore (Security & AI Journals)GHAttack: Generative Adversarial Attacks on Heterogeneous Graph Neural Networks
Jan 13, 2026InfoResearchPeer-reviewedSecurityResearchGHAttack is a generative adversarial attack method against heterogeneous graph neural networks (HGNNs). A trained perturbation generator produces a perturbation for each target node in a single forward pass, modifying edges across heterogeneous relations to degrade predictions on target nodes. The authors report high efficiency and effectiveness across ten HGNNs and six datasets.
IEEE Xplore (Security & AI Journals)Abstract Gradient Training: A Unified Certification Framework for Data Poisoning, Unlearning, and Differential Privacy
Dec 31, 2025InfoResearchPeer-reviewedResearchSecurityAbstract Gradient Training (AGT) is a unified framework for certifying a given model and training procedure against training data perturbations, covering bounded perturbations, removal of data points, and addition of new samples. It works by bounding the reachable set of parameters to establish provable parameter-space bounds, and it targets models trained with first-order optimization methods. The framework is presented as covering data poisoning, machine unlearning, and differential privacy.
JMLR (Journal of Machine Learning Research)On the Relevance of Byzantine Robust Optimization Against Data Poisoning
Dec 31, 2025InfoResearchPeer-reviewedResearchSecurityThe paper examines whether Byzantine machine learning, where distributed workers may deviate arbitrarily from the algorithm, is relevant to data poisoning, a weaker threat model where only local datasets are corrupted. The authors prove that Byzantine-robust schemes yield optimal solutions under both fully-poisonous and partially-poisonous local data, and that fully-poisonous workers are more harmful when local data is heterogeneous.
JMLR (Journal of Machine Learning Research)Enhanced Masking-Differential Prompting (MDP): Defending Backdoor Attacks for Pre-trained Language Models Under Few-Shot Learning
Dec 22, 2025InfoResearchPeer-reviewedSecurityResearchThis research paper proposes two enhancements to masking-differential prompting (MDP), a defense against backdoor attacks on pre-trained language models under few-shot, prompt-based learning. The authors replace KL divergence with Jensen–Shannon divergence, which stays finite when anchor set density is low, and add an adaptive threshold method that searches automatically using a false rejection rate (FRR) allowance instead of costly manual ROC/AUC threshold selection. The authors report that the method defends better against typical backdoor attacks on text classification and generation benchmarks.
IEEE Xplore (Security & AI Journals)Trap: Mitigating Poisoning-Based Backdoor Attacks by Treating Poison With Poison
Dec 15, 2025InfoResearchPeer-reviewedSecurityResearchResearchers propose a training method that detects poisoned samples in the early stages of training and removes the backdoor by retraining the classifier part of the model on relabeled poisoned samples. Evaluated against twelve attacks on four datasets, it reduced the average attack success rate to 0.07% while lowering average accuracy by 0.33%.
Fix: The proposed defense detects poisoned samples early in training and removes the backdoor by retraining the classifier part of the model on relabeled poisoned samples.
IEEE Xplore (Security & AI Journals)PPFPL: Cross-Silo Privacy-Preserving Federated Prototype Learning Against Data Poisoning Attacks
Dec 12, 2025InfoResearchPeer-reviewedSecurityResearchThis research article proposes PPFPL, a privacy-preserving federated prototype learning framework for cross-silo federated learning. It uses prototypes as client-submitted model updates to limit the effect of poisoned Non-IID data, and a secure aggregation protocol built on homomorphic encryption that runs on two servers to achieve Byzantine-robust aggregation. The authors report theoretical convergence and privacy analyses and experiments on public datasets showing resistance to data poisoning under Non-IID settings.
IEEE Xplore (Security & AI Journals)Why Not Diversify Triggers? APK-Specific Backdoor Attack Against Android Malware Detection
Dec 11, 2025LowResearchPeer-reviewedSecurityResearchThis paper presents APK-Specific Backdoor Attack (ASBA), a poisoning method against machine learning-based Android malware detection (AMD) models. ASBA trains a generative adversarial network to create a distinct trigger for each malware sample, so that discovering one trigger via static analysis does not expose all the malware carrying it. The authors report a 94.6% average attack success rate across three datasets, five feature extraction methods and three classification models.
IEEE Xplore (Security & AI Journals)Enhancing the Security of Large Character Set CAPTCHAs Using Transferable Adversarial Examples
Dec 9, 2025InfoResearchPeer-reviewedResearchSecurityResearchers propose ACG (Adversarial Large Character Set CAPTCHA Generation), a framework for defending large character set CAPTCHAs against automated attacks. It pairs a Fine-grained Generation Module with an Ensemble Generation Module that adds global perturbations, and they release Adv-Eval, a toolkit with datasets from 10 popular Chinese CAPTCHA schemes. Experiments show the average success rate of diverse attacks falls from 51.52% to 2.56%.
IEEE Xplore (Security & AI Journals)
Topic added 2026-10-09. An item belongs to this topic when its title matches one of the topic's patterns or its summary mentions the topic at least twice. Report a wrong match with the feedback button on the item.