Adversarial machine learning
Attacks on how models learn and decide: adversarial examples, evasion, data poisoning and backdoors in trained models.
- All items
- 108
- Last 90 days
- 47
- Change
- +114%vs 22 before
Items per month
| Month | Items |
|---|---|
| May 2025 | 0 |
| Jun 2025 | 0 |
| Jul 2025 | 0 |
| Aug 2025 | 1 |
| Sep 2025 | 1 |
| Oct 2025 | 1 |
| Nov 2025 | 1 |
| Dec 2025 | 6 |
| Jan 2026 | 4 |
| Feb 2026 | 6 |
| Mar 2026 | 11 |
| Apr 2026 | 8 |
| May 2026 | 10 |
| Jun 2026 | 9 |
| Jul 2026 | 8 |
| Aug 2026 | 3 |
| Sep 2026 | 24 |
| Oct 2026 | 12 |
101 items
Exploring backdoor attack and defense algorithms in LLMS: Enhancing in-context learning security
Sep 28, 2026InfoResearchPeer-reviewedSecurityResearchThis paper shows that an attacker can manipulate LLM behavior by poisoning the demonstration context used in in-context learning, without fine-tuning the model. The authors present ICLAttack, a backdoor method that poisons demonstration examples or demonstration prompts, reporting a 95.0% average attack success rate on OPT models across three datasets. They also propose ICLDefense, which uses a lightweight auxiliary model and an ensemble-based strategy to refine LLM outputs and reduce attack success.
Fix: ICLDefense: a defense algorithm that utilizes model ensembles, employing a lightweight auxiliary model to refine LLM outputs through an ensemble-based strategy, which the source says substantially reduces the attack success rate compared to existing methods while preserving model performance.
OpenAlex (peer-reviewed AI security)Detection of Adversarial Attacks on Super-Resolvers Using Spectral Features
Sep 28, 2026InfoResearchPreprintSecurityResearchResearchers propose a spectral detection method for finding adversarial attacks embedded in super-resolution model weights, a preprocessing component of imaging pipelines. The method trains an XGBoost detector on the radially-averaged power spectral density and benchmarks it against magnitude- and phase-based Fourier spectrum detectors across training and cross-architecture scenarios. The proposed detector outperforms the comparison detectors in most scenarios, and high-frequency features prove most informative for detecting AdvSR attacks.
Arxiv (cs.RO + cs.CV security)IndirectAD: Practical Data Poisoning Attacks Against Recommender Systems for Item Promotion
Sep 28, 2026LowResearchPeer-reviewedSecurityResearchResearchers introduce IndirectAD, a data poisoning attack against recommender systems inspired by Trojan attacks on machine learning. The attack first promotes a trigger item, then transfers that advantage to a target item by creating co-occurrence data between them, which reduces the number of controlled accounts needed. Experiments on multiple datasets and recommender systems show noticeable impact with only 0.05% of a platform's user base.
IEEE Xplore (Security & AI Journals)Stealthy Physical Adversarial Attacks on Speaker Recognition via Near-Ultrasonic Perturbations
Sep 28, 2026LowResearchPeer-reviewedSecurityResearchAdvNup is a physical adversarial attack that spoofs deep neural network speaker recognition systems using near-ultrasonic perturbations played through commercial off-the-shelf speakers, avoiding the specialized hardware that earlier ultrasound attacks required. The authors use single-sideband modulation with low-pass filtering, a nonlinear frequency response model, and time-frequency masking to keep the adversarial signal intact through physical transmission. In simulated and physical experiments, it reached a 100% attack success rate for closed-set identification and over 90% for open-set identification in targeted attacks.
IEEE Xplore (Security & AI Journals)Enhancing the Robustness of Speaker Recognition Systems in Noisy Channels Using Thermal Diffusion-Based Adversarial Attacks and Wavelet-Based Defence
Sep 27, 2026InfoResearchPeer-reviewedSecurityResearchPublished in Circuits Systems and Signal Processing on 2026-09-28, this research paper concerns the robustness of speaker recognition systems in noisy channels. It combines thermal diffusion-based adversarial attacks with a wavelet-based defence. The source text provided contains only the publication details and reference list, so the method and findings are not described.
OpenAlex (peer-reviewed AI security)One Attack to Fool Them All: Highly Transferable Black-Box Adversarial Attacks on Frontier MLLMs
Sep 27, 2026InfoResearchPreprintSecurityResearchResearchers ask whether a single adversarial image can consistently mislead diverse frontier MLLMs in black-box settings. They present O-Attack, a transfer-based black-box framework that exploits cross-modally aligned semantic representations in surrogate models, and report attack success rates rising on GPT-5.4 (29.1% to 77.2%), Claude-4.6 (42.8% to 81.6%), and Gemini-3.1 (38.2% to 80.9%). Across 24 MLLMs, O-Attack outperforms six state-of-the-art methods in black-box transferability.
Arxiv (cs.RO + cs.CV security)Benchmarking post-processing methods in local differential privacy for utility and adversarial robustness
Sep 26, 2026InfoResearchPeer-reviewedResearchPrivacyElsevier Security JournalsAdaptive Defense Optimization Under Intelligent Data Poisoning Attacks in Industrial Control System
Sep 25, 2026InfoResearchPeer-reviewedSecurityResearchThis paper proposes a detection-aware stochastic optimization framework for industrial control systems facing asynchronous, unreliable feedback and adversarial data poisoning. It uses the SWaT dataset to model poisoning behavior and an adaptive acknowledgment (ACK) bundling mechanism, driven by a multi-objective stochastic gradient descent (SGD) algorithm with a residual-based anomaly indicator, to adjust feedback timing. Numerical experiments on an autonomous surface vehicle (ASV) benchmark show closed-loop recovery under AI-driven data poisoning attacks.
Fix: The source describes an adaptive ACK bundling mechanism that regulates feedback timing and adjusts the ACK bundling window online when anomalous behavior is detected, with Lyapunov–Krasovskii (LK)-based conditions established to ensure closed-loop stability. It does not describe a patch, fixed version or configuration fix for a specific product.
IEEE Xplore (Security & AI Journals)Advanced Cross-Attack Backdoor Detector Based on Disturbance Immunity Learned From Classic Backdoor Attacks
Sep 17, 2026InfoResearchPeer-reviewedSecurityResearchThe paper presents the Advanced Cross-attack Backdoor Detector (ACBD), which detects trigger-injected samples by solving a labeled binary classification task based on disturbance immunity, rather than unlabeled feature clustering. ACBD is trained on one class of a poisoned dataset (1/100 of CIFAR-100) with two classic attacks, using a small LSTM with 53 K parameters and at most 10 clean images for perturbation. The authors report state-of-the-art detection with cross-attack generalization, including on unseen triggers and different target labels.
IEEE Xplore (Security & AI Journals)EXE-Bench: Ranking the Tradeoffs of AI-Based Windows Malware Detectors for Real-World Usability
Sep 17, 2026InfoResearchPeer-reviewedSecurityResearchExisting evaluations of AI-based Windows malware detectors differ in training and test data, lack temporal analysis, skip adversarial content-injection tests, and ignore deployment compute costs, so they cannot show which detector to deploy. The authors introduce EXE-Bench, which assesses performance, temporal and adversarial robustness, and computational overhead, combining them into one score for direct comparison. Their analysis finds that feature-engineered domain knowledge remains highly useful, resisting both time and adversarial attacks, while most deep networks excel only right after deployment.
IEEE Xplore (Security & AI Journals)CAFBA: Context-aware adaptive fusion backdoor attack for polyp segmentation
Sep 9, 2026InfoResearchPeer-reviewedSecurityResearchElsevier Security JournalsDefending Against Adversarial Malware Attacks on ML-Based Android Malware Detection Methods
Sep 2, 2026InfoResearchPeer-reviewedSecurityResearchResearchers propose ADD, a plug-in defense framework that improves the adversarial robustness of ML-based Android malware detection (AMD) methods against problem space attacks, which generate real adversarial malware rather than only adversarial feature vectors. The authors report that ADD performs well against the evaluated state-of-the-art problem space attacks across multiple ML-based AMD methods. They also state that ADD enhances the adversarial robustness of real-world antivirus solutions.
Fix: ADD, a plug-in adversarial Android malware defense framework that enhances the adversarial robustness of ML-based AMD methods against problem space attacks.
IEEE Xplore (Security & AI Journals)Action-Level Backdoor Attacks Against Deep Reinforcement Learning Systems via Adaptive Reward Exploration
Sep 2, 2026InfoResearchPeer-reviewedSecurityResearchAdapdoor is an attack framework that injects action-level backdoors into deep reinforcement learning models during training, so that adversaries can manipulate action outputs at deployment. It initializes the backdoor reward from benign reward statistics and iteratively fine-tunes it using performance feedback from benign and backdoor tasks. Across 3 DRL algorithms, 11 environments and 53 backdoor tasks, it outperforms existing baselines by 42.0% to 144.2%.
Fix: The source says the authors evaluate three potential defenses to explore pathways for mitigating this threat, but it does not name them or state a fix.
IEEE Xplore (Security & AI Journals)Generative Textual Adversarial Attack Through Extensible Compositional Perturbation via Reinforcement Learning for Policy Optimization
Sep 1, 2026InfoResearchPeer-reviewedSecurityResearchGECOMP is a generative textual adversarial attack that uses reinforcement learning for policy optimization. An LLM-based generator rewrites input sequences using an extensible library of compositional perturbations, constrained by semantic similarity and edit magnitude. Across four public datasets and five victim models, it achieved higher attack success rates than ten baseline methods while using fewer queries and less edit magnitude.
IEEE Xplore (Security & AI Journals)WPEBA: A Novel Ensemble Black-Box Adversarial Attack for Visual Recognition Systems via Wavelet Packet Decomposition
Sep 1, 2026InfoResearchPeer-reviewedSecurityResearchWPEBA is a frequency-driven ensemble black-box adversarial attack for visual recognition systems that uses wavelet packet decomposition to split images into frequency sub-bands. It adjusts sub-band weights from internal gradient feedback and updates surrogate-model weights from target-model query feedback, reaching an average attack success rate of nearly 99% with one or two queries across six standard architectures. The authors report it also remains effective against defended models and the Google Cloud Vision API.
IEEE Xplore (Security & AI Journals)3DGAA: Realistic and Robust 3D Gaussian-Based Adversarial Attack for Autonomous Driving
Sep 1, 2026InfoResearchPeer-reviewedSecurityResearchResearchers present 3DGAA, a framework for physical adversarial attacks on camera-based perception in autonomous vehicles. It uses 3D multi-view optimization with a 3D Gaussian splatting surrogate to produce print-only vehicle wraps that keep geometry unchanged. In CARLA simulations and miniature-vehicle tests, the wraps substantially reduced detection confidence and average precision across multiple modern detectors.
IEEE Xplore (Security & AI Journals)Edge-Only Universal Adversarial Attacks in Distributed Learning
Aug 19, 2026InfoResearchPeer-reviewedSecurityResearchResearchers study whether universal adversarial perturbations (UAPs) can be generated when an attacker controls only the edge portion of a split model, meaning its initial network layers. They introduce edge-only untargeted and targeted UAP formulations that manipulate intermediate features before the split point. On ImageNet, the attacks transfer strongly to the unknown cloud component and compare favorably with classical white-box and black-box techniques.
IEEE Xplore (Security & AI Journals)Adversarial Purification by Consistency-Aware Latent Space Optimization on Data Manifolds
Aug 14, 2026InfoResearchPeer-reviewedSecurityResearchThis paper proposes Consistency Model-based Adversarial Purification (CMAP), which removes adversarial perturbations from inputs to deep neural networks by optimizing vectors in the latent space of a pre-trained consistency model. The method combines a perceptual consistency restoration mechanism, a latent distribution consistency constraint, and an ensemble-based latent vector consistency prediction scheme. Experiments on CIFAR-10 and ImageNet-100 report significantly improved robustness against strong adversarial attacks while preserving high natural accuracy.
IEEE Xplore (Security & AI Journals)Text Adversarial Attacks With Dynamic Outputs
Aug 12, 2026InfoResearchPeer-reviewedSecurityResearchResearchers introduce the Textual Dynamic Outputs Attack (TDOA), a text adversarial attack for settings where the number and content of model labels vary, such as LLM outputs and multi-label classification. TDOA trains a clustering-based surrogate model that approximates dynamic fine-grained outputs with static coarse-grained labels, and uses a farthest-label targeted strategy to induce larger output changes. Evaluated on five datasets and ten victim models including GPT-4o and GPT-4.1, it reaches an 80.8% maximum attack success rate with five queries per text.
IEEE Xplore (Security & AI Journals)A comprehensive analysis of adversarial attacks against spam filters
Jul 25, 2026InfoResearchPeer-reviewedSecurityResearchElsevier Security Journals
Topic added 2026-10-09. An item belongs to this topic when its title matches one of the topic's patterns or its summary mentions the topic at least twice. Report a wrong match with the feedback button on the item.