aisecwatch.com
DashboardVulnerabilitiesNewsResearchArchiveStatsDatasetFor devs
Subscribe
aisecwatch.com

Real-time AI security monitoring. Tracking AI-related vulnerabilities, safety and security incidents, privacy risks, research developments, and policy changes.

Navigation

VulnerabilitiesNewsResearchDigest ArchiveNewsletter ArchiveSubscribeData SourcesStatisticsDatasetAPIIntegrationsWidgetRSS Feed

Maintained by

Truong (Jack) Luu

Information Systems Researcher

Research

Academic papers, new techniques, benchmarks, and theoretical findings in AI/LLM security.

to
Export CSV
1217 items

v5.1.0

inforesearchIndustry
securityresearch
Nov 6, 2025

ATLAS Data v5.1.0 is an updated framework that documents security threats and defenses related to AI systems, now containing 16 tactics, 84 techniques, and 32 mitigations. The update adds new attack methods targeting AI, such as prompt injection (tricking an AI by hiding instructions in its input), deepfake generation, and data theft from AI services, along with new defensive measures like human oversight of AI agent actions and restricted permissions for AI tools. It also includes 42 real-world case studies showing how these attacks and defenses apply in practice.

MITRE ATLAS Releases

FUBA: Backdoor Federated Learning via Federated Unlearning

inforesearchPeer-Reviewed
security

Source Camera Model Identification via Federated Learning Using Laplacian-Based Patches

inforesearchPeer-Reviewed
research

MaxDiv: Zero-Shot Machine Unlearning via Distributionally Divergent Erasing Samples

inforesearchPeer-Reviewed
research

A Systematic Literature Review on SWOT Analysis of Prompt Engineering Techniques

inforesearchPeer-Reviewed
research

Lightweight Reparameterizable Integral Neural Networks for Mobile Applications

inforesearchPeer-Reviewed
research

Co-AttenDWG: Coattentive Dimension-Wise Gating and Expert Fusion for Multimodal Offensive Content Detection

inforesearchPeer-Reviewed
research

Federated Learning of Dynamic Bayesian Network via Continuous Optimization From Time Series Data

inforesearchPeer-Reviewed
research

v5.0.1

inforesearchIndustry
industry

v5.0.0

inforesearchIndustry
security

Maritime Operational Technology Shipboard Testbed for Cybersecurity Research, Training, Exercises, and Education

inforesearchPeer-Reviewed
security

Asynchronous Federated Learning With Nonconvex Client Objective Functions and Heterogeneous Dataset

inforesearchPeer-Reviewed
research

A Mathematical Certification for Positivity Conditions in Neural Networks With Applications to Partial Monotonicity and Trustworthy AI

inforesearchPeer-Reviewed
research

Action-Perturbation Backdoor Attacks on Partially Observable Multiagent Systems

inforesearchPeer-Reviewed
security

Privacy Protection of Dual Averaging Push for Decentralized Optimization via Zero-Sum Structured Perturbations

inforesearchPeer-Reviewed
research

Do More With Less: Architecture-Agnostic and Data-Free Extraction Attack Against Tabular Model

inforesearchPeer-Reviewed
security

Advancing Biometric Authentication With Dual-Threshold Multi-Modal Systems and Geometric Programming for Enhanced Digital Security

inforesearchPeer-Reviewed
security

Really Unlearned? Verifying Machine Unlearning via Influential Sample Pairs

inforesearchPeer-Reviewed
security

Engineering Trustworthy AI: A Developer Guide for Empirical Risk Minimization

inforesearchPeer-Reviewed
research

A Deep Reinforcement Learning Approach to Time Delay Differential Game Deception Resource Deployment

inforesearchPeer-Reviewed
research
Previous57 / 61Next
research
Nov 6, 2025

Researchers discovered a new attack called FUBA (federated unlearning backdoor attack) that exploits a privacy feature in federated learning (a technique where multiple parties train an AI model together without sharing their raw data). The attack uses malicious unlearning requests, which are supposed to let participants remove their data from a trained model, to secretly inject backdoors (hidden harmful behaviors) into the model instead. The attack is difficult to detect because it hides from existing security defenses.

IEEE Xplore (Security & AI Journals)
Nov 5, 2025

This research proposes FedFFTNet, a system for identifying which camera model took a photo by using federated learning (a technique where AI models train on data kept private across multiple devices rather than sharing raw data centrally). The system uses a lightweight deep learning architecture and a Laplacian-based patch selection strategy (focusing on sharp, detailed areas of images) to identify cameras while maintaining privacy, achieving very high accuracy rates on standard benchmark datasets.

IEEE Xplore (Security & AI Journals)
privacy
Oct 30, 2025

This article presents MaxDiv, a technique for machine unlearning, which is the process of removing specific knowledge from an AI model after training to protect privacy, even when the original training data is no longer available. MaxDiv works by creating special synthetic data samples that have opposite characteristics to the data being forgotten, and it uses knowledge distillation (a technique where a model learns to replicate another model's behavior) to ensure important information isn't accidentally lost during the unlearning process.

IEEE Xplore (Security & AI Journals)
Oct 28, 2025

This article reviews prompt engineering (the practice of designing inputs like questions or instructions to guide AI systems toward better responses) and analyzes its strengths, weaknesses, opportunities, and threats using a SWOT framework. The review covers how prompt engineering can improve interactions with large language models (advanced AI systems trained on vast amounts of text) across industries like healthcare and education, while also identifying challenges around maintaining accuracy and efficiency.

IEEE Xplore (Security & AI Journals)
Oct 27, 2025

This paper presents RINNs (reparameterizable integral neural networks), a new type of AI model designed to run efficiently on mobile devices with limited computing power. The key innovation is a reparameterization strategy that converts the complex mathematical structure used during training into a simpler feed-forward structure (a straightforward sequence of processing steps) at inference time, allowing these models to achieve high accuracy (79.1%) while running very fast (0.87 milliseconds) on mobile hardware.

IEEE Xplore (Security & AI Journals)
Oct 20, 2025

This paper presents Co-AttenDWG, a new method for detecting offensive content by combining text and images together. The approach uses coattention (a technique where two types of data pay attention to each other simultaneously), dimension-wise gating (a mechanism that selectively emphasizes important features at a detailed level), and expert fusion (combining predictions from multiple specialized models) to better understand how text and visual information relate to each other.

IEEE Xplore (Security & AI Journals)
Oct 16, 2025

This research presents a federated learning (FL, a technique where multiple parties train an AI model together without sharing raw data) approach for learning the structure of dynamic Bayesian networks (DBN, a statistical model that represents relationships between variables over time) from distributed time series data. The method addresses challenges like data privacy and heterogeneity (when different parties' data follows different patterns), and provides mathematical proof that the approach reliably converges to good results, which hadn't been formally guaranteed before in this setting.

IEEE Xplore (Security & AI Journals)
Oct 15, 2025

N/A -- The provided content is a navigation menu and feature listing from GitHub's website, not a security issue, vulnerability report, or technical problem related to AI/LLMs.

MITRE ATLAS Releases
research
Oct 15, 2025

ATLAS Data v5.0.0 introduces a new "Technique Maturity" field that categorizes AI attack techniques based on evidence level, ranging from feasible (proven in research) to realized (used in actual attacks). The release adds 11 new techniques covering AI agent attacks like context poisoning (injecting false information into an AI system's memory), credential theft from AI configurations, and prompt injection (tricking an AI by hiding malicious instructions in its input), plus updates to existing techniques and case studies.

MITRE ATLAS Releases
Oct 15, 2025

This article describes a maritime operational technology shipboard testbed, which is a controlled testing environment on a ship that mimics real maritime systems. The testbed allows cybersecurity researchers and professionals to safely study cyberattacks and test defensive strategies without risking actual ships or critical systems.

IEEE Xplore (Security & AI Journals)
Oct 14, 2025

This research addresses challenges in asynchronous federated learning (AFL, a distributed machine learning approach where multiple devices train a model on their own data without sending raw data to a central server), specifically when devices have different types of objective functions and uneven data. The authors propose two main improvements: a staleness-aware aggregation mechanism (a method that reduces the influence of outdated updates from slower devices) and a dynamic learning rate schedule (an adaptive parameter that adjusts training speed based on how delayed each device's updates are) to improve model accuracy and stability in real-world environments where devices have different computing power and network speeds.

Fix: The source explicitly proposes two solutions: (1) 'a staleness-aware aggregation mechanism that penalizes outdated updates, ensuring fresher data have a more significant influence on the global model,' and (2) 'a dynamic learning rate schedule that adapts to client staleness and heterogeneity, improving stability and convergence.' The authors demonstrate practical implementation using 'PyTorch and Python's asyncio library.'

IEEE Xplore (Security & AI Journals)
safety
Oct 14, 2025

This research presents LipVor, an algorithm that mathematically verifies whether a trained neural network (a computer model with interconnected nodes that learns patterns) follows partial monotonicity constraints, which means outputs change predictably with certain inputs. The method works by testing the network at specific points and using mathematical properties to guarantee the network behaves correctly across its entire domain, potentially allowing neural networks to be used in critical applications like credit scoring where trustworthiness and predictable behavior are required.

IEEE Xplore (Security & AI Journals)
research
Oct 13, 2025

Researchers discovered a type of backdoor attack (hidden malicious instructions planted in AI systems) on multiagent reinforcement learning systems, where one adversary agent uses its actions to trigger hidden failures in other agents' decision-making policies. Unlike previous attacks that assumed unrealistic direct control over what victims observe, this attack is more practical because it works through normal agent interactions in partially observable environments (where agents cannot always see what others are doing). The researchers developed a training method to help adversary agents efficiently trigger these backdoors with minimal suspicious actions.

IEEE Xplore (Security & AI Journals)
privacy
Oct 13, 2025

This research addresses privacy risks in decentralized optimization (where multiple networked computers work together to solve a problem without a central coordinator) by proposing ZS-DDAPush, an algorithm that adds mathematical noise structures to protect sensitive node information during communication. The key innovation is that ZS-DDAPush achieves privacy protection while maintaining the accuracy and efficiency of the optimization process, avoiding the typical trade-offs seen in other privacy methods like differential privacy (adding statistical noise to protect individual data) or encryption (scrambling data so only authorized parties can read it).

IEEE Xplore (Security & AI Journals)
research
Oct 13, 2025

Researchers developed TabExtractor, a tool that can steal tabular models (AI systems trained on spreadsheet-like data) without needing access to the original training data or knowing how the model was built. The attack works by creating synthetic data samples and using a special neural network architecture called a contrastive tabular transformer (CTT, a type of AI that learns by comparing similar and different examples) to reverse-engineer a clone of the victim model that performs almost as well as the original. This research shows that tabular models face serious security risks from extraction attacks.

IEEE Xplore (Security & AI Journals)
Oct 13, 2025

This article describes BMMA-GPT, a biometric authentication system that uses multiple forms of identification (like fingerprints and facial recognition) together with mathematical optimization to improve security and speed. The system uses a dual-threshold approach (two decision points to verify identity) and can be tailored to different organizational needs, achieving high accuracy while keeping verification time under 1.5 seconds.

IEEE Xplore (Security & AI Journals)
research
Oct 13, 2025

Machine unlearning allows AI models to forget the effects of specific training samples, but verifying whether this actually happened is difficult because existing checks (like backdoor attacks or membership inference attacks, which test if a model remembers data by trying to extract or manipulate it) can be fooled by a dishonest model provider who simply retrains the model to pass the test rather than truly unlearning. This paper proposes IndirectVerify, a formal verification method that uses pairs of connected samples (trigger samples that are unlearned and reaction samples that should be affected by that unlearning) with intentional perturbations (small changes to training data) to create indirect evidence that unlearning actually occurred, making it harder to fake.

IEEE Xplore (Security & AI Journals)
safety
Oct 13, 2025

AI systems used for important decisions often rely on empirical risk minimization (ERM, a training method that reduces prediction errors on known data) to build models, but these systems can suffer from unintentional bias, lack of transparency, and other risks. The EU has established Ethics Guidelines requiring trustworthy AI to meet seven key requirements, yet current ERM-based design prioritizes accuracy over trustworthiness. This article argues that developers need to balance four core objectives when designing AI systems: fairness (not discriminating against groups), privacy (protecting user data), robustness (resisting intentional attacks like fake news), and explainability (being transparent about how decisions are made).

IEEE Xplore (Security & AI Journals)
security
Oct 10, 2025

This research proposes a new method for deploying cyber deception (defensive tricks to confuse attackers) in networks by combining deep reinforcement learning (a type of AI that learns by trial and error) with game theory that accounts for time delays. The method uses an algorithm called proximal policy optimization (PPO, a technique for training AI to make optimal decisions) to figure out where and when to place deception resources, and tests show it outperforms existing approaches in handling complex network attacks.

IEEE Xplore (Security & AI Journals)