Adversarial machine learning
Attacks on how models learn and decide: adversarial examples, evasion, data poisoning and backdoors in trained models.
- All items
- 108
- Last 90 days
- 47
- Change
- +114%vs 22 before
Items per month
| Month | Items |
|---|---|
| May 2025 | 0 |
| Jun 2025 | 0 |
| Jul 2025 | 0 |
| Aug 2025 | 1 |
| Sep 2025 | 1 |
| Oct 2025 | 1 |
| Nov 2025 | 1 |
| Dec 2025 | 6 |
| Jan 2026 | 4 |
| Feb 2026 | 6 |
| Mar 2026 | 11 |
| Apr 2026 | 8 |
| May 2026 | 10 |
| Jun 2026 | 9 |
| Jul 2026 | 8 |
| Aug 2026 | 3 |
| Sep 2026 | 24 |
| Oct 2026 | 12 |
101 items
AdvDiffusion: Adversarial Patches Generation for Face Recognition With High Transferability in Physical Domain
Mar 2, 2026InfoResearchPeer-reviewedSecurityResearchThe paper presents AdvDiffusion, a method for generating adversarial face patches that can be pasted onto a face to cause face recognition models to make false identity judgments. It first selects a patch region using facial gradient maps, then adds noise to an image and denoises it with a pre-trained diffusion model, guided by an adversarial loss. Experiments report effective attacks on black-box face recognition models in both digital and physical settings, with better robustness to complex physical environments than some state-of-the-art methods.
IEEE Xplore (Security & AI Journals)Complementary Text-Guided Attention for Zero-Shot Adversarial Robustness
Mar 2, 2026InfoResearchPeer-reviewedSecurityResearchThe authors examine how adversarial perturbations shift text-guided attention in CLIP, a pre-trained vision-language model. They propose TGA-ZSR, which uses a Local Attention Refinement Module and a Global Attention Constraint Module, and Comp-TGA, which adds complementary class-prompt and non-class-prompt foreground attention. Across 16 datasets, TGA-ZSR and Comp-TGA improve zero-shot robust accuracy over the prior state of the art by 9.58% and 11.95%, respectively.
IEEE Xplore (Security & AI Journals)PPOM-Attack: A Substitute Model-Free Perturbation Prediction and Optimization Method for Black-Box Adversarial Attack Against Face Recognition
Feb 23, 2026LowResearchPeer-reviewedSecurityResearchPPOM-Attack is a black-box attack against face recognition that avoids substitute models by taking feedback directly from the target model. It uses a proximal policy optimization (PPO)-based agent to predict and disturb perturbation regions in a face image, and a minimum brightness offsets method to keep adversarial images high quality. The authors report a 21.7% average gain in attack success rate over state-of-the-art FR attacks across seven FR models.
IEEE Xplore (Security & AI Journals)AdvScan: Black-Box Adversarial Example Detection at Runtime Through Power Analysis
Feb 9, 2026InfoResearchPeer-reviewedSecurityResearchAdvScan is a black-box method for detecting adversarial examples on TinyML models at runtime by analyzing power consumption. It builds a baseline of power signatures from benign inputs and applies a one-sample t-test to flag inputs whose signatures deviate significantly. Evaluated on three MLPerf Tiny models across two STM32 microcontrollers and 318,400 test inputs, it detected 99.984% of adversarial examples, with 40 false negatives and no false positives.
IEEE Xplore (Security & AI Journals)Practical and Flexible Backdoor Attack Against Deep Learning Models via Shell Code Injection
Feb 9, 2026InfoResearchPeer-reviewedSecurityResearchResearchers propose shell code injection (SCI), a non-poisoning backdoor attack framework against deep learning models that needs no training. It uses a logic-driven backdoor shell with trigger consistency verification and short-circuit code packaging to force wrong predictions, and an LLM-assisted module generates attack target code adaptively. Experiments report almost 100% attack success rate (ASR) across various settings.
IEEE Xplore (Security & AI Journals)PROTheft: A Projector-Based Model Extraction Attack in the Physical World
Feb 6, 2026InfoResearchPeer-reviewedSecurityResearchPROTheft is a model extraction attack that extends digital-domain attacks to vision-based devices in the physical world. An attacker uses a projector placed in front of the on-board camera to feed attack samples to a black-box target, and a simulation module compensates for detail loss in the digital-to-physical-to-digital transformation. On a public autonomous driving dataset, the authors report over 80% fidelity with the target model and an mAP 50 above 0.85.
IEEE Xplore (Security & AI Journals)Allies Teach Better Than Enemies: Inverse Adversaries for Robust Knowledge Distillation
Feb 3, 2026InfoResearchPeer-reviewedResearchSecurityThis paper proposes an adversarially robust knowledge distillation method that compresses a large robust teacher model into a lightweight student. It refines inputs into inverse adversarial examples by reversing the sign of the adversarial perturbation, then applies gradient matching between teacher and student on those inputs, plus a weight-space disruption strategy. On ImageNet, the authors report roughly 3.8% gains over prior methods in both clean and robust accuracy.
IEEE Xplore (Security & AI Journals)A Wolf in Sheep’s Clothing: Unveiling a Stealthy Backdoor Attack in Subgraph Federated Learning
Jan 28, 2026InfoResearchPeer-reviewedSecurityResearchResearchers propose BEEF, an end-to-end backdoor attack against subgraph federated learning for node classification, where subgraphs are distributed across devices. BEEF uses a trigger generator trained jointly with the backdoored model, crafting adversarial perturbations as triggers that cause misclassification without changing model parameters. Evaluations across eight datasets, four models, five attacks and six aggregation methods report effectiveness against GNNs with minimal impact on normal data performance.
IEEE Xplore (Security & AI Journals)GHAttack: Generative Adversarial Attacks on Heterogeneous Graph Neural Networks
Jan 13, 2026InfoResearchPeer-reviewedSecurityResearchGHAttack is a generative adversarial attack method against heterogeneous graph neural networks (HGNNs). A trained perturbation generator produces a perturbation for each target node in a single forward pass, modifying edges across heterogeneous relations to degrade predictions on target nodes. The authors report high efficiency and effectiveness across ten HGNNs and six datasets.
IEEE Xplore (Security & AI Journals)Abstract Gradient Training: A Unified Certification Framework for Data Poisoning, Unlearning, and Differential Privacy
Dec 31, 2025InfoResearchPeer-reviewedResearchSecurityAbstract Gradient Training (AGT) is a unified framework for certifying a given model and training procedure against training data perturbations, covering bounded perturbations, removal of data points, and addition of new samples. It works by bounding the reachable set of parameters to establish provable parameter-space bounds, and it targets models trained with first-order optimization methods. The framework is presented as covering data poisoning, machine unlearning, and differential privacy.
JMLR (Journal of Machine Learning Research)On the Relevance of Byzantine Robust Optimization Against Data Poisoning
Dec 31, 2025InfoResearchPeer-reviewedResearchSecurityThe paper examines whether Byzantine machine learning, where distributed workers may deviate arbitrarily from the algorithm, is relevant to data poisoning, a weaker threat model where only local datasets are corrupted. The authors prove that Byzantine-robust schemes yield optimal solutions under both fully-poisonous and partially-poisonous local data, and that fully-poisonous workers are more harmful when local data is heterogeneous.
JMLR (Journal of Machine Learning Research)Enhanced Masking-Differential Prompting (MDP): Defending Backdoor Attacks for Pre-trained Language Models Under Few-Shot Learning
Dec 22, 2025InfoResearchPeer-reviewedSecurityResearchThis research paper proposes two enhancements to masking-differential prompting (MDP), a defense against backdoor attacks on pre-trained language models under few-shot, prompt-based learning. The authors replace KL divergence with Jensen–Shannon divergence, which stays finite when anchor set density is low, and add an adaptive threshold method that searches automatically using a false rejection rate (FRR) allowance instead of costly manual ROC/AUC threshold selection. The authors report that the method defends better against typical backdoor attacks on text classification and generation benchmarks.
IEEE Xplore (Security & AI Journals)Trap: Mitigating Poisoning-Based Backdoor Attacks by Treating Poison With Poison
Dec 15, 2025InfoResearchPeer-reviewedSecurityResearchResearchers propose a training method that detects poisoned samples in the early stages of training and removes the backdoor by retraining the classifier part of the model on relabeled poisoned samples. Evaluated against twelve attacks on four datasets, it reduced the average attack success rate to 0.07% while lowering average accuracy by 0.33%.
Fix: The proposed defense detects poisoned samples early in training and removes the backdoor by retraining the classifier part of the model on relabeled poisoned samples.
IEEE Xplore (Security & AI Journals)PPFPL: Cross-Silo Privacy-Preserving Federated Prototype Learning Against Data Poisoning Attacks
Dec 12, 2025InfoResearchPeer-reviewedSecurityResearchThis research article proposes PPFPL, a privacy-preserving federated prototype learning framework for cross-silo federated learning. It uses prototypes as client-submitted model updates to limit the effect of poisoned Non-IID data, and a secure aggregation protocol built on homomorphic encryption that runs on two servers to achieve Byzantine-robust aggregation. The authors report theoretical convergence and privacy analyses and experiments on public datasets showing resistance to data poisoning under Non-IID settings.
IEEE Xplore (Security & AI Journals)Why Not Diversify Triggers? APK-Specific Backdoor Attack Against Android Malware Detection
Dec 11, 2025LowResearchPeer-reviewedSecurityResearchThis paper presents APK-Specific Backdoor Attack (ASBA), a poisoning method against machine learning-based Android malware detection (AMD) models. ASBA trains a generative adversarial network to create a distinct trigger for each malware sample, so that discovering one trigger via static analysis does not expose all the malware carrying it. The authors report a 94.6% average attack success rate across three datasets, five feature extraction methods and three classification models.
IEEE Xplore (Security & AI Journals)Enhancing the Security of Large Character Set CAPTCHAs Using Transferable Adversarial Examples
Dec 9, 2025InfoResearchPeer-reviewedResearchSecurityResearchers propose ACG (Adversarial Large Character Set CAPTCHA Generation), a framework for defending large character set CAPTCHAs against automated attacks. It pairs a Fine-grained Generation Module with an Ensemble Generation Module that adds global perturbations, and they release Adv-Eval, a toolkit with datasets from 10 popular Chinese CAPTCHA schemes. Experiments show the average success rate of diverse attacks falls from 51.52% to 2.56%.
IEEE Xplore (Security & AI Journals)Versatile Backdoor Attack With Visible, Semantic, Sample-Specific and Compatible Triggers
Dec 9, 2025InfoResearchPeer-reviewedSecurityResearchResearchers introduce the Visible, Semantic, Sample-specific, and Compatible (VSSC) trigger for backdoor attacks on deep neural networks. They present an automated pipeline that selects triggers with large language models, inserts them into images with generative models, and checks insertion quality with vision-language models. The authors report that VSSC stays effective under visual distortions and works in physical scenarios using corresponding real-world objects.
IEEE Xplore (Security & AI Journals)Investigating the Robustness of Fuzzy Deep Learning on Noisy Medical Images
Nov 24, 2025InfoResearchPeer-reviewedResearchSecurityThis research tests a deep neuro-fuzzy system (DNFS) for medical image classification under six noise types and adversarial attacks. The DNFS matched state-of-the-art models on original images across seven biomedical datasets and showed superior average accuracy on noisy data. It was also tested for susceptibility to adversarial perturbations, which the authors report exposes weaknesses in deep learning.
IEEE Xplore (Security & AI Journals)Action-Perturbation Backdoor Attacks on Partially Observable Multiagent Systems
Oct 13, 2025InfoResearchPeer-reviewedSecurityResearchResearchers study backdoor attacks on reinforcement learning agents in partially observable multiagent systems. Instead of modifying a victim's observations directly, an adversary agent uses its own actions to affect what other agents observe, and a trained trigger policy, guided by auxiliary rewards, lets it activate their backdoors with few actions. Experiments show the method triggers others' backdoors efficiently, and the authors also study defenses.
IEEE Xplore (Security & AI Journals)AI-Shielder: Exploiting Backdoors to Defend Against Adversarial Attacks
Sep 29, 2025InfoResearchPeer-reviewedSecurityResearchAI-Shielder is a defense that deliberately embeds backdoors in a DNN so that adversarial perturbations fail while the main task keeps working. Evaluated against sixteen adversarial example generation approaches, it reduces the attack success rate from 91.8% to 3.8%, outperforming state-of-the-art works by 37.2%, with a 0.6% decline in clean data accuracy and 1.43% overhead to model prediction time.
IEEE Xplore (Security & AI Journals)
Topic added 2026-10-09. An item belongs to this topic when its title matches one of the topic's patterns or its summary mentions the topic at least twice. Report a wrong match with the feedback button on the item.