Adversarial machine learning
Attacks on how models learn and decide: adversarial examples, evasion, data poisoning and backdoors in trained models.
- All items
- 108
- Last 90 days
- 47
- Change
- +114%vs 22 before
Items per month
| Month | Items |
|---|---|
| May 2025 | 0 |
| Jun 2025 | 0 |
| Jul 2025 | 0 |
| Aug 2025 | 1 |
| Sep 2025 | 1 |
| Oct 2025 | 1 |
| Nov 2025 | 1 |
| Dec 2025 | 6 |
| Jan 2026 | 4 |
| Feb 2026 | 6 |
| Mar 2026 | 11 |
| Apr 2026 | 8 |
| May 2026 | 10 |
| Jun 2026 | 9 |
| Jul 2026 | 8 |
| Aug 2026 | 3 |
| Sep 2026 | 24 |
| Oct 2026 | 12 |
108 items
Beyond Stop Signs: Why Evasion Attacks Matter Even More
May 13, 2026InfoResearchPeer-reviewedSecurityResearchThe source argues that evasion attacks on machine learning systems deserve more concern than their usual treatment suggests. Canonical demonstrations have stayed largely academic for over a decade, and their intellectual appeal has led many to dismiss evasion as having little operational consequence.
IEEE Xplore (Security & AI Journals)CVE-2026-31230: Adversarial Robustness Toolbox code injection via command-line arguments
May 12, 2026HighVulnerabilitySecurityCVE-2026-31230The Adversarial Robustness Toolbox (ART) through 1.20.1 has a command-line argument injection flaw in its Kubeflow component, robustness_evaluation_fgsm_pytorch.py. The script passes the --clip_values and --input_shape arguments to the unsafe eval() function, so injected Python code runs when eval() is called. An attacker who can control these arguments, for example through pipeline configuration or automated scripts, can achieve arbitrary code execution on the system running the ART evaluation.
NVD/CVE DatabaseCVE-2026-31229: The Adversarial Robustness Toolbox (ART) thru 1.20.1 contains an insecure deserialization vulnerability (CWE-502) in…
May 12, 2026HighVulnerabilitySecurityCVE-2026-31229The Adversarial Robustness Toolbox (ART) through 1.20.1 contains an insecure deserialization flaw (CWE-502) in its Kubeflow component's model loading. During robustness evaluation, model weights loaded from a file such as model.pt use torch.load() without weights_only=True, which permits Pickle deserialization of arbitrary Python objects. An attacker who uploads a crafted model file to object storage referenced by the pipeline, or who controls the model_id parameter to point to such a file, can achieve remote code execution when the pipeline loads the model.
NVD/CVE DatabaseCVE-2026-31228: Adversarial Robustness Toolbox remote code execution in Kubeflow via eval()
May 12, 2026CriticalVulnerabilitySecurityCVE-2026-31228CVE-2026-31228 affects the Adversarial Robustness Toolbox (ART) through 1.20.1, in its Kubeflow component. The robustness evaluation function for PyTorch models passes user-supplied strings for the LossFn and Optimizer parameters to eval() without sanitization, so an attacker can supply a crafted string containing arbitrary Python code. That code runs when eval() is called, giving the attacker full control of the system running the ART evaluation.
NVD/CVE DatabaseAdversarial Robustness of Link Sign Prediction in Signed Graphs
May 5, 2026InfoResearchPeer-reviewedSecurityResearchResearchers show that balance theory, which signed graph neural networks (SGNNs) use to model positive and negative relationships, creates exploitable weaknesses under closed-box attacks. They propose balance-attack, which targets graph balance degree and is solved with a heuristic for an NP-hard problem, and introduce Balance Augmented-Signed Graph Contrastive Learning (BA-SGCL) to restore robustness by maintaining high balance degree in the latent space. Experiments across multiple SGNN architectures and real-world datasets report both the attack's effectiveness and BA-SGCL's improved resilience.
Fix: BA-SGCL, a contrastive learning framework with balance augmentation, maintains high balance degree in the latent space to improve model resilience against balance-attack.
IEEE Xplore (Security & AI Journals)UAP4MA: Leveraging Multi-Agent Bandits to Generate Universal Adversarial Perturbations for Malware Attribution
May 4, 2026InfoResearchPeer-reviewedSecurityResearchUAP4MA is a decision-based, problem-space, black-box method for generating universal adversarial perturbations against APT malware attribution models. It uses a multi-agent Multi-Armed Bandit framework that applies functionality-preserving transformations from 13 types. On the AMG43 and APTMalware datasets, it achieves over four times the fooling rate, double the attack success rate, and an 80% reduction in training time compared with state-of-the-art methods.
IEEE Xplore (Security & AI Journals)Optimizing stealthiness in universal adversarial perturbations via class-selective and perceptual similarity metrics
Apr 20, 2026InfoResearchPeer-reviewedSecurityResearchStealthy-UAP is a framework for generating universal adversarial perturbations that are harder to detect. It combines a Class-Selective Attack, which targets only semantically consistent classes and keeps predictions for non-target inputs unchanged, with Perceptual Similarity Optimization, which uses high-level deep features to align perturbations with the Human Visual System. The authors report that it outperforms state-of-the-art baselines on ImageNet subsets and the full ImageNet validation set, and remains effective under mainstream adversarial defenses.
Elsevier Security JournalsBioGuard: Malicious sample free defense method for biometric classifiers against model extraction attacks
Apr 17, 2026InfoResearchPeer-reviewedSecurityResearchThe source is a Computers & Security article by Ziting Ren, Yucong Duan and Qi Qi, published online on 16 April 2026, titled "BioGuard: Malicious sample free defense method for biometric classifiers against model extraction attacks." The provided text contains only the title, publication date, source and authors, with no abstract or findings.
Elsevier Security JournalsDefending Against Patch-Based and Texture-Based Adversarial Attacks With Spectral Decomposition
Apr 16, 2026InfoResearchPeer-reviewedSecurityResearchResearchers propose Adversarial Spectrum Defense (ASD), which uses Discrete Wavelet Transform (DWT) spectral decomposition to analyze adversarial patterns across multiple frequency scales. Combined with off-the-shelf Adversarial Training (AT), ASD+AT reportedly achieves state-of-the-art performance against patch-based and texture-based attacks, outperforming previous defenses by 21.73% in the AP metric, including against adaptive adversaries designed against ASD.
IEEE Xplore (Security & AI Journals)Transferable Adversarial Attack on Referring Video Object Segmentation
Apr 13, 2026InfoResearchPeer-reviewedSecurityResearchThis work presents the first comprehensive study of adversarial robustness in referring video object segmentation (RVOS) models, which segment text-referred objects in video and matter for safety-critical uses such as autonomous driving. The authors propose xM-ICM, a cross-prompt multimodal attack that jointly corrupts visual and textual embeddings and uses two momentum banks to keep perturbations coherent across clips. Experiments on three benchmarks and five RVOS models report strong white-box attack performance and good black-box transferability.
IEEE Xplore (Security & AI Journals)HENet: A Heterogeneous Encoding Network for General and Robust Adversarial Example Generation
Apr 13, 2026InfoResearchPeer-reviewedSecurityResearchHENet, a dual-branch generative model, is proposed to build a robust generator-based adversarial attack framework against deep neural networks. Its Adaptive Feature Fusion Module addresses dimension and representativeness mismatches between CNNs and Transformers, while a Dynamic Differentiable JPEG Simulator uses an adaptive quantization mask to improve robustness under JPEG compression. The authors report a better attack success rate, lower perturbation magnitude and higher robustness across target architectures.
IEEE Xplore (Security & AI Journals)ReSLC: Defending backdoor attacks on intelligent vulnerability detection via redundant semantic LLM compression
Apr 8, 2026InfoResearchPeer-reviewedResearchSecurityResearchers from Shanghai Jiao Tong University and collaborators (Hongying Zhang and others) published ReSLC in the Journal of Information Security and Applications, Volume 100, in July 2026. The source text provided contains only publication metadata (date, journal, volume, authors) and no description of the method, findings or numbers.
Elsevier Security JournalsAISM: Adversarial image steganography model for defending unauthorized recognition
Apr 3, 2026InfoResearchPeer-reviewedResearchSecurityElsevier Security JournalsSeeking Flat Minima Over Diverse Surrogates for Improved Adversarial Transferability: A Theoretical Framework and Algorithmic Instantiation
Apr 1, 2026InfoResearchPeer-reviewedSecurityResearchThis paper presents a theoretical framework for transfer-based black-box adversarial attacks, deriving a transferability bound that links adversarial transferability to flat minima over a surrogate model set and the adversarial model discrepancy. Based on this bound, the authors build a surrogate model set with diverse adversarial vulnerabilities and generate a model-Diversity-compatible Reverse Adversarial Perturbation (DRAP). Experiments on the NIPS2017 and CIFAR-10 datasets against various target models show the proposed attack is effective.
IEEE Xplore (Security & AI Journals)One Trigger, Multiple Victims: Clean-Label Neighborhood Backdoor Attacks on Graph Neural Networks
Mar 27, 2026InfoResearchPeer-reviewedSecurityResearchThis paper proposes a clean-label backdoor attack on Graph Neural Networks that injects a single trigger node attached to a target node. The attack misclassifies the target and its one-hop neighbors as the target class without changing their features or labels. Across multiple real-world benchmarks and mainstream GNN architectures, the authors report over 95% attack success rate on target nodes and their neighbors in most settings, including under state-of-the-art defenses.
IEEE Xplore (Security & AI Journals)PadNet: Defending Neural Networks Against Adversarial Examples
Mar 25, 2026InfoResearchPeer-reviewedSecurityResearchACM Digital Library (TOPS, DTRAP, CSUR)Assessing and Improving DNN Robustness Against Adversarial Examples From the Perspective of Fully Connected Layers
Mar 25, 2026InfoResearchPeer-reviewedSecurityResearchResearchers propose a redundant fully connected layer that can be plugged into existing image classification model backbones to improve adversarial robustness. The layer is trained with a loss function that uses cosine similarity to maximize the difference and diversity among multiple fully connected parts. Tested against 12 representative white-box and black-box attacks on two popular datasets, the approach reportedly gives significant robustness with negligible extra training overhead and hardly degrades clean-sample accuracy.
Fix: The proposed mitigation is the redundant fully connected layer, combined with existing model backbones in a pluggable manner, trained with the cosine-similarity-based loss function described in the source.
IEEE Xplore (Security & AI Journals)Filter, Obstruct, and Dilute: Defending Against Backdoor Attacks on Semi-Supervised Learning
Mar 25, 2026InfoResearchPeer-reviewedSecurityResearchThis research paper analyzes how backdoor attacks work against semi-supervised learning (SSL), finding that attackers exploit pseudo-labeling to build stronger trigger-target correlations and that triggers can appear in low, medium, or high frequency bands. The authors propose Backdoor Invalidator (BI), a defense combining complementary learning, trigger mix-up, and dual domain filtering. Evaluation against state-of-the-art attacks shows BI significantly reduces average attack success rate while keeping comparable clean-data accuracy.
Fix: BI, a defense framework integrating complementary learning, trigger mix-up, and dual domain filtering, which the authors describe as a plug-in component. Code is available at https://github.com/wxr99/Backdoor_Invalidator4SSL
IEEE Xplore (Security & AI Journals)SRAP: Robust and Transferable Self-Reversible Adversarial Patch for Image Privacy Protection
Mar 23, 2026InfoResearchPeer-reviewedSecurityResearchThe authors introduce SRAP, a Self-Reversible Adversarial Patch technique for generating reversible adversarial examples to protect images from malicious deep model identification and analysis. Combining small, non-overlapping adversarial patches with prediction error expansion data hiding, the method reports transferability of up to 90% or higher between models and an 88% attack success rate on commercial APIs.
IEEE Xplore (Security & AI Journals)A Dual-Purpose Framework for Backdoor Defense and Backdoor Amplification in Diffusion Models
Mar 18, 2026InfoResearchPeer-reviewedSecurityResearchPureDiffusion is a dual-purpose framework for diffusion models that inverts backdoor triggers for defense and reinforces them for attack amplification. Its defense uses two novel loss functions based on trigger-induced distribution shifts across timesteps and denoising consistency, followed by a detection method. Reported results include near-perfect detection accuracy and attack success rates of nearly 100% for existing backdoor attacks, with training time reduced by up to 20x.
IEEE Xplore (Security & AI Journals)
Topic added 2026-10-09. An item belongs to this topic when its title matches one of the topic's patterns or its summary mentions the topic at least twice. Report a wrong match with the feedback button on the item.