aisecwatch.com
DashboardVulnerabilitiesNewsResearchArchiveStatsDatasetFor devs
Subscribe
aisecwatch.com

Real-time AI security monitoring. Tracking AI-related vulnerabilities, safety and security incidents, privacy risks, research developments, and policy changes.

Navigation

VulnerabilitiesNewsResearchDigest ArchiveNewsletter ArchiveSubscribeData SourcesStatisticsDatasetAPIIntegrationsWidgetRSS Feed

Maintained by

Truong (Jack) Luu

Information Systems Researcher

Research

Academic papers, new techniques, benchmarks, and theoretical findings in AI/LLM security.

to
Export CSV
1217 items

Practical Continual Forgetting for Pre-Trained Vision Models

inforesearchPeer-Reviewed
researchprivacy
Jan 15, 2026

This research addresses how to remove unwanted information from pre-trained vision models (AI systems trained to understand images) when users or model owners request it, especially when these deletion requests come one after another over time. The researchers propose Group Sparse LoRA (GS-LoRA), a technique that uses Low-Rank Adaptation modules (efficient add-on components that modify specific neural network layers) to selectively forget targeted classes or information while keeping the rest of the model working well, even when some training data is missing.

Fix: The paper proposes two explicit solutions: (1) Group Sparse LoRA (GS-LoRA), which uses Low-Rank Adaptation modules to fine-tune Feed-Forward Network layers in Transformer blocks for each forgetting task independently, combined with group sparse regularization to automatically select and zero out specific LoRA groups. (2) GS-LoRA++, an extension that incorporates prototype information as additional supervision, moving logits (output scores) away from the original prototype of forgotten classes while pulling logits closer to prototypes of remaining classes.

IEEE Xplore (Security & AI Journals)

BlindU: Blind Machine Unlearning Without Revealing Erasing Data

inforesearchPeer-Reviewed
research

Robust Physics-Based Deep MRI Reconstruction via Diffusion Purification

inforesearchPeer-Reviewed
research

SLeak: Multi-Target Privacy Stealing Attack Against Split Learning

inforesearchPeer-Reviewed
security

GHAttack: Generative Adversarial Attacks on Heterogeneous Graph Neural Networks

inforesearchPeer-Reviewed
security

Armor: Shielding Unlearnable Examples Against Data Augmentation

inforesearchPeer-Reviewed
security

Model Lineage Analysis: Determination and Closeness Measurement

inforesearchPeer-Reviewed
research

Boosting Adversarial Training With Mitigating Hard Sample Interference

inforesearchPeer-Reviewed
research

$\ell_{0}$-Norm Penalty Embedded Feature Selection in Universum Learning

inforesearchPeer-Reviewed
research

DeepFilter: A Transformer-Style Framework for Accurate and Efficient Process Monitoring

inforesearchPeer-Reviewed
research

Dopamine Neuron Activity Visualization From MRI Sensors Through 3-D MRI-to-SPECT Synthesis for Parkinson’s Disease Assessment

inforesearchPeer-Reviewed
research

Accurate and Robust Neural Architecture Search via a Flexible Supernet

inforesearchPeer-Reviewed
research

Revisiting Out-of-Distribution Detection in Real-Time Object Detection: From Benchmark Pitfalls to a New Mitigation Paradigm

inforesearchPeer-Reviewed
research

Reinforcement Learning-Based Optimal Formation Tracking for UAVs With Safety Constraints

inforesearchPeer-Reviewed
research

The Sample Complexity of Parameter-Free Stochastic Convex Optimization

inforesearchPeer-Reviewed
research

Probabilistic Rainfall Downscaling: Joint Generalized Neural Models with Censored Spatial Gaussian Copula

inforesearchPeer-Reviewed
research

A Unified Approach to Analysis and Design of Denoising Markov Models

inforesearchPeer-Reviewed
research

Enhancing Accuracy in Generative Models via Knowledge Transfer

inforesearchPeer-Reviewed
research

Error Analyses of Auto-Regressive Video Diffusion Models

inforesearchPeer-Reviewed
research

Nonparametric Estimation of a Factorizable Density using Diffusion Models

inforesearchPeer-Reviewed
research
Previous50 / 61Next
privacy
Jan 15, 2026

BlindU is a method that allows users to remove their data's influence from trained AI models while keeping that data hidden from the server. Instead of uploading raw data to the server (which creates privacy risks), BlindU lets users create compressed versions of their data locally, and the server performs the removal process only on these compressed versions, making it practical for federated learning (a distributed training setup where data stays on users' devices).

Fix: BlindU implements unlearning through several stated mechanisms: (1) 'the user locally generates privacy-preserving representations, and the server performs unlearning solely on these representations and their labels', (2) use of an information bottleneck mechanism that 'learns representations that distort maximum task-irrelevant information from inputs', (3) 'two dedicated unlearning modules tailored explicitly for IB-based models and uses a multiple gradient descent algorithm to balance forgetting and utility retaining', and (4) 'a noise-free differential privacy masking method to deal with the raw erasing data before compressing' for additional privacy protection.

IEEE Xplore (Security & AI Journals)
safety
Jan 14, 2026

Deep learning models used for MRI reconstruction (creating medical images from incomplete data) can fail when faced with unexpected situations like noise, different imaging settings, or unseen medical conditions. This paper proposes RODIO, a method that uses diffusion models (AI systems that gradually refine noisy data into clear images) as "purifiers" to make MRI reconstruction systems more reliable, and shows it works better than existing robustification techniques like adversarial training (deliberately exposing models to bad inputs during training to make them stronger).

Fix: The paper proposes RODIO as the solution: using pretrained diffusion models as purifiers to improve robustness by fine-tuning on purified examples, which eliminates the need for adversarial training's complex optimization process. The authors state their approach demonstrates adaptability across multiple deep learning MRI reconstruction models, compatibility with accelerated diffusion samplers, robustness to data with unseen lesions, and effectiveness with unsupervised generative reconstructors.

IEEE Xplore (Security & AI Journals)
research
Jan 14, 2026

Split Learning (SL) is a distributed learning framework designed to preserve privacy while reducing computational load, but researchers discovered a new attack called SLeak that allows a server adversary to steal client data and models. The attack works by exploiting information in the smashed data (intermediate data passed between clients and server) and server model to build a substitute client that mimics the target client's behavior, without needing strong privacy assumptions or much auxiliary data. The study shows SLeak is more effective than previous attacks across different datasets and scenarios.

IEEE Xplore (Security & AI Journals)
research
Jan 13, 2026

This research paper introduces GHAttack, a new method for attacking heterogeneous graph neural networks, or HGNNs (AI systems that learn from complex data structures with multiple types of relationships). Instead of slowly computing attacks through complicated math problems, GHAttack uses a generative model (an AI trained to create outputs) to quickly generate perturbations (small modifications) that damage how well HGNNs make predictions on target nodes. The authors tested their method on multiple HGNNs and datasets to show it works efficiently and effectively.

IEEE Xplore (Security & AI Journals)
privacy
Jan 12, 2026

Unlearnable examples are protective noises added to private data to prevent AI models from learning useful information from them, but this paper shows that data augmentation (a common technique that creates variations of training data to improve model performance) can undo this protection and restore learnability from 21.3% to 66.1% accuracy. The researchers propose Armor, a defense framework that adds protective noise while accounting for data augmentation effects, using a surrogate model (a practice model used to simulate the real training process) and smart augmentation selection to keep private data unlearnable even after augmentation is applied.

Fix: The paper proposes Armor, a defense framework that works by: (1) designing a non-local module-assisted surrogate model to better capture the effect of data augmentation, (2) using a surrogate augmentation selection strategy that maximizes distribution alignment between augmented and non-augmented samples to choose the optimal augmentation strategy for each class, and (3) using a dynamic step size adjustment algorithm to enhance the defensive noise generation process. The authors state that 'Armor can preserve the unlearnability of protected private data under data augmentation' and plan to open-source the code upon publication.

IEEE Xplore (Security & AI Journals)
Jan 12, 2026

This research addresses how to identify whether one machine learning model is derived from another model through modification techniques (adjusting or fine-tuning an existing model rather than training from scratch), and how to measure how much two models differ from each other. The authors propose a method that determines lineage (derivative relationships) by checking if two models' parameters exist in the same local optimum of the loss landscape (the mathematical space of possible model configurations), and measure closeness by analyzing how their decision boundaries (the lines or surfaces that separate different predictions) differ from each other.

IEEE Xplore (Security & AI Journals)
safety
Jan 6, 2026

This research addresses a problem in adversarial training (a technique that teaches AI models to resist adversarial examples, which are inputs carefully designed to fool the model). When adversarial training tries to improve both normal accuracy and robustness at the same time, it struggles with hard samples (data points near the decision boundary where the model finds it difficult to classify correctly), often forcing a sacrifice of one goal for the other. The authors propose MHSI (mitigating hard sample interference), which uses two techniques: a weighted adaptive mechanism that helps the model focus more on learning clean samples, and a dynamic calibration strategy guided by gradient analysis that adjusts how the model handles hard samples, resulting in improved robustness without losing accuracy.

IEEE Xplore (Security & AI Journals)
Jan 5, 2026

This research paper describes a new machine learning method that combines Universum learning (using unlabeled data from the same domain as labeled training data to improve model training) with feature selection (choosing the most important input variables). The authors add an L0-norm penalty (a mathematical constraint that forces the model to use fewer features) to a support vector machine classifier, and develop an algorithm to solve this optimization problem efficiently.

IEEE Xplore (Security & AI Journals)
Jan 5, 2026

DeepFilter is a modified AI framework based on Transformers (a type of neural network architecture) designed to monitor industrial processes more accurately and efficiently. Standard Transformers use self-attention (a mechanism where the model weighs the importance of different parts of input data), but this approach struggles with process monitoring because it doesn't capture meaningful patterns in logs and requires a lot of computation. DeepFilter replaces the self-attention layer with an efficient filtering layer that better identifies long-term patterns while using less computing power.

IEEE Xplore (Security & AI Journals)
Jan 5, 2026

This article describes a medical imaging framework that converts MRI (magnetic resonance imaging, a non-invasive scanning technique) scans into synthetic SPECT images (single photon emission computed tomography, a radiological imaging technique that typically requires radioactive drug injections) to help detect Parkinson's disease by analyzing dopamine neuron activity. The proposed 3-D-cycle conversion network uses a machine learning approach inspired by CycleGAN (a type of neural network that learns to transform images from one style to another) to generate SPECT-like images from regular MRI data without needing radioactive injections. The researchers claim their method outperforms similar existing approaches in accuracy.

IEEE Xplore (Security & AI Journals)
safety
Jan 5, 2026

Neural architecture search (NAS, the automated process of designing AI model structures) can produce models that are vulnerable to adversarial attacks (manipulated inputs designed to fool AI systems). This paper presents ARNAS++, a method that searches for neural architectures that are both accurate and robust against adversarial attacks by using a flexible supernet (a large parent network from which smaller networks are derived) with adjustable parameter budgets and width.

IEEE Xplore (Security & AI Journals)
safety
Jan 5, 2026

Out-of-distribution (OoD, inputs that don't match what an AI was trained on) detection in object detection systems causes AI models to make overconfident wrong predictions on objects they shouldn't recognize. This paper reveals that popular benchmark datasets used to test OoD detection have quality problems, where up to 13% of test objects are mislabeled, making current methods appear better than they really are. The authors propose a new training-time approach where object detectors are fine-tuned using carefully created OoD training data that looks similar to normal objects, which reduces false detections by 91% in YOLO models.

Fix: The paper introduces a training-time mitigation paradigm where 'we fine-tune the detector using a carefully synthesized OoD dataset that semantically resembles in-distribution objects.' This approach 'shapes a defensive decision boundary by suppressing objectness on OoD objects' and achieves 'a 91% reduction in hallucination error of a YOLO model on BDD-100 K.' The methodology is shown to work across multiple detection architectures including YOLO, Faster R-CNN, and RT-DETR.

IEEE Xplore (Security & AI Journals)
Jan 1, 2026

This article presents a control method for multiple fixed-wing UAVs (unmanned aerial vehicles, or drones) that need to fly together in formation while avoiding collisions and handling unpredictable disturbances. The approach uses reinforcement learning (a type of AI that learns by trial and error) combined with control barrier functions (mathematical tools that enforce safety constraints) to create a system that keeps the UAVs safe and stable while optimizing their performance.

IEEE Xplore (Security & AI Journals)
Dec 31, 2025

This research addresses how stochastic convex optimization (a machine learning technique for finding the best solution by processing data in random batches) can work when key problem parameters are unknown. The authors propose two methods: a model selection technique that prevents overfitting (when an AI learns noise in the validation data instead of real patterns), and a regularization-based approach that estimates unknown parameters to achieve optimal efficiency. Experiments on image classification and shape-counting tasks show these methods help reduce overfitting on small validation sets.

JMLR (Journal of Machine Learning Research)
Dec 31, 2025

This research presents a method for converting large-scale weather predictions into detailed local rainfall forecasts using neural networks and statistical models. The approach works in two steps: first, it uses joint generalized neural models (neural networks that predict the parameters of probability distributions) to estimate rainfall distributions based on coarse weather data, and second, it uses a censored latent Gaussian copula (a mathematical model that captures how rainfall patterns are related across nearby locations) to ensure spatial coherence. The method was tested on UK weather data and performed better than existing techniques.

JMLR (Journal of Machine Learning Research)
Dec 31, 2025

This paper presents a mathematical framework for understanding denoising Markov models (generative models that learn to reverse a noising process to create new data). The authors use concepts from statistical mechanics to establish rigorous rules for how these models work, unifying existing approaches like diffusion models and proposing new variations using different types of mathematical processes.

JMLR (Journal of Machine Learning Research)
Dec 31, 2025

This paper studies how to improve the accuracy of generative models (AI systems that create new data, like images or text) by using knowledge transfer, where a model trained on one task helps train a model on a different task. The researchers introduce a framework based on 'Shared Embedding,' a technique that finds common patterns between different tasks even when their data looks different, and show that this approach improves performance in two types of generative models: diffusion models (which gradually refine random noise into structured outputs) and normalizing flows (mathematical transformations that learn data distributions).

JMLR (Journal of Machine Learning Research)
Dec 31, 2025

Auto-regressive video diffusion models (AR-VDMs, systems that generate videos by predicting one frame at a time) struggle with two problems: history forgetting, where they lose track of earlier frames they created, and temporal degradation, where video quality gets worse over time. Researchers created Meta-ARVDM, a theoretical framework that analyzes both problems and shows that using more past frames reduces history forgetting, while also introducing a new way to evaluate these models using a "needle-in-a-haystack" test (finding specific content buried in larger data).

JMLR (Journal of Machine Learning Research)
Dec 31, 2025

This research paper studies diffusion models, a type of AI used to generate images and audio, as a statistical method for density estimation (learning the probability distribution of data). The authors show that when data has a factorizable structure (meaning it can be broken into independent low-dimensional components, like in Bayesian networks), diffusion models can efficiently learn this structure and achieve optimal performance using a specially designed sparse neural network architecture (one where most connections between neurons are inactive).

JMLR (Journal of Machine Learning Research)