Adversarial machine learning
Attacks on how models learn and decide: adversarial examples, evasion, data poisoning and backdoors in trained models.
- All items
- 108
- Last 90 days
- 47
- Change
- +114%vs 22 before
Items per month
| Month | Items |
|---|---|
| May 2025 | 0 |
| Jun 2025 | 0 |
| Jul 2025 | 0 |
| Aug 2025 | 1 |
| Sep 2025 | 1 |
| Oct 2025 | 1 |
| Nov 2025 | 1 |
| Dec 2025 | 6 |
| Jan 2026 | 4 |
| Feb 2026 | 6 |
| Mar 2026 | 11 |
| Apr 2026 | 8 |
| May 2026 | 10 |
| Jun 2026 | 9 |
| Jul 2026 | 8 |
| Aug 2026 | 3 |
| Sep 2026 | 24 |
| Oct 2026 | 12 |
101 items
PREFed: An Effective and Stealthy Static-Anchor Backdoor Attack via Trigger Pre-Optimization in Federated Learning
Jul 23, 2026InfoResearchPeer-reviewedSecurityResearchPREFed is a static-anchor backdoor attack for federated learning that pre-optimizes trigger patterns on clean data before training, removing the need for round-wise adaptation. The authors report that it reaches over 80% backdoor accuracy within five communication rounds while cutting main task accuracy by less than 2%, versus more than 15% degradation in prior methods, across six defenses on image benchmarks and SST-2.
IEEE Xplore (Security & AI Journals)NEO: Navigating Entropy in Optimized Closed-Box Video Adversarial Attacks
Jul 22, 2026InfoResearchPeer-reviewedSecurityResearchNEO is a closed-box adversarial attack on deep learning video recognition systems that uses information-entropy guidance to cut query costs. The authors model query-based attacks as an open-box approximation and find that queries near decision boundaries carry the most information. NEO combines diffusion-based perturbation initialization with mutual-information-driven optimization and reports significantly higher efficiency and stealthiness than existing methods on four benchmark datasets.
IEEE Xplore (Security & AI Journals)Rethinking Fake Adversarial Examples for Single-Step Adversarial Training
Jul 20, 2026InfoResearchPeer-reviewedSecurityResearchResearchers examine why single-step adversarial training, a cheaper alternative to standard adversarial training, often suffers catastrophic overfitting under larger perturbations. They identify "fakers," single-step adversarial examples that are learned and correctly classified yet fail to expose true model vulnerabilities, and that degrade robustness, resist learning and diverge strongly from their clean counterparts. They propose FAST, which adjusts label smoothing by learning difficulty and adds a weak-effect auxiliary sample, reporting superior clean accuracy and robustness across various attacks.
IEEE Xplore (Security & AI Journals)Spa: Stealthy and Persistent Backdoor Attacks in Federated Learning via Feature-Space Alignment
Jul 17, 2026InfoResearchPeer-reviewedSecurityResearchResearchers propose Spa, a backdoor attack framework for federated learning that aims to be both stealthy and persistent. Instead of training a conflicting secondary task, Spa uses feature-space alignment to fold backdoor features into the primary objective, and it uses adversarial dynamic trigger optimization that co-evolves with the global model. Experiments report attack success rates near 100% with minimal utility loss, and the backdoor stays effective around 900 FL rounds after attacks stop.
IEEE Xplore (Security & AI Journals)PANDA: Diffusion-Guided Purification and Adaptation for Robust Point Cloud Classification Against Adversarial Attack
Jul 17, 2026InfoResearchPeer-reviewedSecurityResearchDeep learning point cloud classifiers are highly vulnerable to adversarial attacks, and existing diffusion-based purification defenses have a distributional gap and a semantic mismatch. The authors propose PANDA, a two-stage framework: PANDA-P trains a dual-branch diffusion purifier on both clean-to-clean and adversarial-to-clean paths, and PANDA-A fine-tunes the classifier with a consistency-driven objective to recalibrate its decision boundary. The source reports consistently superior robustness over existing purification-based defenses on synthetic and real-world benchmarks.
IEEE Xplore (Security & AI Journals)Perspective-Invariant Attack With Enhanced Transferability of Adversarial Examples
Jul 16, 2026InfoResearchPeer-reviewedSecurityResearchThis research paper proposes the Perspective-Invariant Attack (PIA), which improves the transferability of adversarial examples between black-box DNN models. PIA uses a multi-DOF vertex sampling strategy covering perspective transformations from 2-DOF translation to 8-DOF projective mapping. The authors also present PIA-Mix, an extension that combines this transformation with auxiliary methods, and report that both outperform state-of-the-art transfer-based attacks across DNN architectures, defenses and multimodal LLMs.
IEEE Xplore (Security & AI Journals)Beyond Single-Pair Attacks: Disrupting Vision-Language Pre-Training Models With Dual-Semantic Frequency Stealth
Jul 14, 2026InfoResearchPeer-reviewedSecurityResearchResearchers propose DSFG-Attack, a transferable adversarial attack against Vision-Language Pre-training (VLP) models that disrupts cross-modal alignment by injecting non-aligned semantics into image-text pairs and restricting perturbations to high-frequency domains for stealth. On Flickr30K image-text retrieval, with TCL as the source model and CLIP_ViT as the target, it raises the black-box attack success rate by an average of 5.08%. The adversarial examples also transfer to captioning and to Multimodal Large Language Models such as GPT-4o and Qwen 2.
IEEE Xplore (Security & AI Journals)Detection and Mitigation Data Poisoning Attacks in Multimodal Online Federated Learning
Jun 25, 2026InfoResearchPeer-reviewedSecurityResearchThis paper presents the first systematic study of data poisoning attacks in Multimodal Online Federated Learning (MMO-FL), a setting where IoT edge devices train models in a decentralized, real-time way across multiple modalities. The authors give a theoretical analysis of how such attacks degrade learning performance, then propose a detection and mitigation algorithm for MMO-FL systems. Experiments on the UCI-HAR and USC-HAD multimodal datasets show the approach detects and mitigates the attacks.
Fix: The source proposes a novel detection and mitigation algorithm tailored to MMO-FL systems, but does not state a specific fix, patch, configuration change or workaround.
IEEE Xplore (Security & AI Journals)NOAE: Noise-Optimized Adversarial Examples for Multivariate Time Series Anomaly Detection of the Industrial Internet of Things
Jun 19, 2026LowResearchPeer-reviewedSecurityResearchThe paper proposes Noise-Optimized Adversarial Examples (NOAE), an attack method against unsupervised deep-learning multivariate time series anomaly detection models used in the Industrial Internet of Things. Experiments show NOAE significantly reduces detection precision on industrial benchmarks, including a 71.51% precision drop on the SMAP dataset at a 0.01 noise amplitude. The authors also propose a Hybrid Adversarial Defense (HAD) training approach that uses data-end random segment replacement augmentation to improve robustness.
Fix: The authors propose Hybrid Adversarial Defense (HAD), a training approach that uses adversarial examples and data-end random segments replacement augmentation to alleviate the weak robustness of anomaly detection models.
IEEE Xplore (Security & AI Journals)MicroPatch: Directed Backdoor Erasing via Victim Parameter Decoupling
Jun 18, 2026InfoResearchPeer-reviewedSecurityResearchMicroPatch is a backdoor erasing approach for deep neural networks that targets the model parameters influenced by backdoor triggers. It reconstructs trigger patterns through reverse engineering, uses influence functions to separate victim parameter components from clean ones, and patches those components. The authors report that across four datasets and four representative backdoor attacks, plus spatial-frequency and frequency-domain attacks, MicroPatch reduces attack success rates more than existing methods while keeping high classification accuracy.
Fix: The source describes MicroPatch itself as the proposed mitigation: patching decoupled victim parameter components to purify the model. No separate fix, patched version or configuration change is stated.
IEEE Xplore (Security & AI Journals)Enhancing Targeted Adversarial Attacks on Large Vision-Language Models via Intermediate Projector
Jun 10, 2026InfoResearchPeer-reviewedSecurityResearchResearchers show that black-box targeted attacks on Large Vision-Language Models can be made more precise by working through the projector, a semantic bridge between vision and language. They propose Intermediate Projector Guided Attack (IPGA), which aligns Q-Former query outputs with a target and transfers across models, and IPGA with Residual Query Alignment (IPGA-R), which also preserves non-target content for fine-grained edits. The authors report that IPGA beats baselines on global attacks, IPGA-R wins on fine-grained attacks, and the method transfers to Google Gemini and OpenAI GPT.
IEEE Xplore (Security & AI Journals)SOOM: A Schedule-Search-Based Operator Obfuscation Method Against Model Extraction Attacks
Jun 10, 2026InfoResearchPeer-reviewedSecurityResearchSOOM is a schedule-search-based operator obfuscation method that defends compiled DNN models against model extraction attacks on standard CPU and GPU backends. Built on TVM, it uses a security-aware learned cost model based on XGBoost gradient boosted trees to balance security and performance. Tests over 105 operator configurations and more than 30,000 tensor computation test cases raised the operator inference failure rate against state-of-the-art extraction attacks to as high as 89%, with performance gains of up to approximately 25.4% in selected cases.
Fix: SOOM: a schedule-search-based operator obfuscation method built on TVM, using a security-aware learned cost model based on XGBoost gradient boosted trees to generate obfuscated executable code for deep learning operators.
IEEE Xplore (Security & AI Journals)SemAlign-PFL:Exploring stealthy and persistent backdoor attacks against personalized federated learning
Jun 6, 2026LowResearchPeer-reviewedSecurityResearchElsevier Security JournalsTrigger as Entity: Backdoor Attacks to Graph-Based Retrieval-Augmented Generation of Large Language Models
Jun 2, 2026InfoResearchPeer-reviewedSecurityResearchResearchers present the first backdoor attacks against graph-based Retrieval-Augmented Generation (RAG) systems used with LLMs. The attacker poisons a crafted corpus in the external database so that trigger entities are inserted into the knowledge graph, causing the model to give attacker-chosen answers only for trigger-containing queries while answering other queries correctly. The authors evaluate three trigger types (word-level, topic-level and semantic-level) with increasing stealth across multiple knowledge databases and language models, and warn of risks to chatbots and agents built on such systems.
IEEE Xplore (Security & AI Journals)BadBone: Backdoor Attacks Against Backbone Models in Visual Prompt Learning
Jun 1, 2026InfoResearchPeer-reviewedSecurityResearchBadBone is a backdoor attack against prompt learning that uses bi-level optimization to compromise a backbone model, so only downstream tasks employing prompt learning inherit the vulnerability. Experiments on three models and three datasets show the targeted and untargeted backdoored models reach high attack performance while keeping utility on pre-training and downstream tasks. The authors report that six model-level defenses, including Neural Cleanse, ABS, MNTD, NAD, CLP and D-BR, are largely ineffective against these models.
IEEE Xplore (Security & AI Journals)mmGuard: A Countermeasure Against Physical Adversarial Attacks on mmWave Radar Sensing
Jun 1, 2026InfoResearchPeer-reviewedSecurityResearchThis paper presents mmGuard, a physics-based defense against physical adversarial attacks (PAAs) on millimeter-wave radar used in autonomous driving and security checking. It detects attacks by identifying spatial phase discontinuities, anomalous radar cross-section patterns, and violations of physico-kinematic relationships, and the authors report over 90% detection accuracy on mmAD, a dataset of over 110,000 annotated radar frames.
Fix: mmGuard performs per-object attack detection and mitigation, compatible with automotive radar update rates, and uses few-shot adaptation to calibrate to unseen settings.
IEEE Xplore (Security & AI Journals)Robustness of Prompting: Enhancing Robustness of Large Language Models Against Prompt Attacks
May 27, 2026InfoResearchPeer-reviewedResearchSecurityResearchers propose robustness of prompting (RoP), a prompting strategy meant to make large language models less sensitive to input perturbations such as typographical errors and slight character order errors. RoP has two stages: Error Correction, which generates adversarial examples and prompts that fix input errors automatically, and Guidance, which builds an optimal guidance prompt from the corrected input. Experiments on arithmetic, commonsense, and logical reasoning tasks show RoP significantly improves robustness against adversarial perturbations with only minimal accuracy degradation compared to clean input.
IEEE Xplore (Security & AI Journals)MDV: Resolving the Auxiliary Data Dilemma in Model Extraction Defenses
May 20, 2026InfoResearchPeer-reviewedSecurityResearchResearchers present Model Defense Variational Autoencoder (MDV), a method against model extraction attacks (MEA), where attackers build a locally trained clone of a deep learning model. MDV replaces realistic auxiliary data, which is often absent, unstable in effect, and leaves some categories less protected, with virtual auxiliary data. The authors say experiments show it addresses these three problems.
IEEE Xplore (Security & AI Journals)Trigger Without Trace: Toward Stealthy Backdoor Attack on Text-to-Image Diffusion Models
May 20, 2026InfoResearchPeer-reviewedSecurityResearchResearchers propose Trigger without Trace (TwT), a stealthy backdoor attack on text-to-image diffusion models. It uses syntactic structures as triggers to break semantic consistency and applies a Kernel Maximum Mean Discrepancy (KMMD) regularizer to align cross-attention distributions with benign samples. The method reports a 97.5% attack success rate and over 98% of backdoor samples bypassing three state-of-the-art detection mechanisms.
IEEE Xplore (Security & AI Journals)A Study of the Removability of Speaker-Adversarial Perturbations
May 13, 2026InfoResearchPeer-reviewedSecurityResearchThis study examines whether speaker-adversarial perturbations, which mislead speaker recognition models, can be removed from speech entirely. The authors test three scenarios of attacker knowledge (ignorant, semi-informed, well-informed) across optimization-based and feedforward perturbation generators on the LibriSpeech dataset. They find perturbations are not eliminated in the ignorant or semi-informed cases, while feedforward perturbations are considerably reduced in the semi-informed case and nearly eliminated in the well-informed case, allowing restoration of the original speech.
IEEE Xplore (Security & AI Journals)
Topic added 2026-10-09. An item belongs to this topic when its title matches one of the topic's patterns or its summary mentions the topic at least twice. Report a wrong match with the feedback button on the item.