aisecwatch.com
DashboardVulnerabilitiesNewsResearchArchiveStatsDatasetFor devs
Subscribe
aisecwatch.com

Real-time AI security monitoring. Tracking AI-related vulnerabilities, safety and security incidents, privacy risks, research developments, and policy changes.

Navigation

VulnerabilitiesNewsResearchDigest ArchiveNewsletter ArchiveSubscribeData SourcesStatisticsDatasetAPIIntegrationsWidgetRSS Feed

Maintained by

Truong (Jack) Luu

Information Systems Researcher

Research

Academic papers, new techniques, benchmarks, and theoretical findings in AI/LLM security.

to
Export CSV
1217 items

HiveTEE: Scalable and Fine-Grained Isolated Domains With RME and MTE Co-Assisted

inforesearchPeer-Reviewed
security
Feb 3, 2026

HiveTEE is a security architecture that divides applications running inside a TEE (Trusted Execution Environment, a secure zone on a processor that protects sensitive operations from the main operating system) into smaller isolated domains, so that if one part is compromised, the damage doesn't spread to the rest. It uses RME (Realm Management Extension, a hardware feature that creates isolated execution spaces) and MTE (Memory Tagging Extension, a feature that prevents certain memory attacks), and testing shows it adds minimal slowdown (less than 3%) to applications.

IEEE Xplore (Security & AI Journals)

Secure Acceleration of Aggregation Queries Over Homomorphically Encrypted Databases

inforesearchPeer-Reviewed
research

Byzantine-Robust and Communication-Efficient Distributed Learning via Compressed Momentum Filtering

inforesearchPeer-Reviewed
research

Allies Teach Better Than Enemies: Inverse Adversaries for Robust Knowledge Distillation

inforesearchPeer-Reviewed
research

Toward Real-World Holistic Privacy-Preserving Person Re-Identification

inforesearchPeer-Reviewed
security

Evaluating and Mitigating Relationship Hallucinations in Large Vision-Language Models

inforesearchPeer-Reviewed
research

EvTexture++: Event-Driven Texture Enhancement for Video Super-Resolution

inforesearchPeer-Reviewed
research

Jailbreak and Guard Aligned Language Models With Only Few In-Context Demonstrations

inforesearchPeer-Reviewed
security

v5.3.0

inforesearchIndustry
industry

v5.2.0

inforesearchIndustry
security

Forgotten Memories

inforesearchPeer-Reviewed
privacy

CA–CI: Integrating Contextual Integrity and the Capabilities Approach for Dignity Considerations in AI Governance

inforesearchPeer-Reviewed
policy

Understanding the Adversarial Landscape of Large Language Models Through the Lens of Attack Objectives

inforesearchPeer-Reviewed
security

Building Trustworthy AI Agents

inforesearchPeer-Reviewed
safety

NAP-Tuning: Neural Augmented Prompt Tuning for Adversarially Robust Vision-Language Models

inforesearchPeer-Reviewed
safety

DriftTrace: Combating Concept Drift in Security Applications Through Detection and Explanation

inforesearchPeer-Reviewed
research

Safeguarding Federated Learning From Data Reconstruction Attacks via Gradient Dropout

inforesearchPeer-Reviewed
research

Model-Driven Learning-Based Physical Layer Authentication for Mobile Wi-Fi Devices

inforesearchPeer-Reviewed
research

SEGA: A Transferable Signed Ensemble Gaussian Black-Box Attack Against No-Reference Image Quality Assessment Models

inforesearchPeer-Reviewed
security

Privacy-Preserving Model Transcription With Differentially Private Synthetic Distillation

inforesearchPeer-Reviewed
research
Previous48 / 61Next
Feb 3, 2026

This research proposes AHEDB (Accelerated Homomorphically Encrypted DataBase), a system designed to speed up database queries on encrypted data using Fully Homomorphic Encryption, or FHE (a method that lets computers perform calculations on encrypted information without decrypting it first). The system uses Encrypted Multiple Maps to reduce computational strain and a Single Range Cover algorithm for indexing, achieving better performance than existing FHE-based approaches while maintaining security.

IEEE Xplore (Security & AI Journals)
Feb 3, 2026

This research addresses two major challenges in distributed learning (training AI models across multiple computers with separate data): Byzantine robustness (protecting against computers that send corrupted or malicious information) and communication efficiency (reducing the amount of data sent between computers). The authors propose a new method using Polyak Momentum (a technique that smooths out noisy updates) to handle both compression of data being sent and attacks from faulty computers, and they prove their approach works better than existing methods.

IEEE Xplore (Security & AI Journals)
safety
Feb 3, 2026

This research proposes a new method for knowledge distillation (training a smaller AI model to mimic a larger one) that preserves adversarial robustness (the ability to resist attacks designed to fool AI systems). Instead of having the student model copy all predictions from the teacher model, the method uses "inverse adversarial examples" (inputs created by reversing the direction of adversarial attacks) to guide learning toward more reliable predictions, resulting in better robustness transfer between models.

IEEE Xplore (Security & AI Journals)
privacy
Feb 3, 2026

Person re-identification (Re-ID, systems that recognize and track individuals across camera footage) systems can be attacked to steal pedestrian images and the AI model itself, threatening privacy for both the system operator and people being monitored. Existing privacy-protection methods fail to defend against all types of leaks while keeping the system working normally, so researchers propose SHIELD, a two-stage framework that uses protected image generation and feature protection techniques to prevent data and model theft without reducing the system's accuracy for authorized users.

IEEE Xplore (Security & AI Journals)
safety
Feb 3, 2026

Large vision-language models (LVMs, AI systems that process both images and text) often make mistakes by hallucinating incorrect relationships between objects in images, such as falsely claiming one object is near another. Researchers created R-Bench, a benchmark (a standardized test) to evaluate these relationship hallucination errors, and found that these mistakes happen because models rely too much on language patterns rather than actually analyzing the visual content. The study proposes Region-Aware Alignment Mitigation (RA²M), which improves the model's attention to specific regions of an image to better align its descriptions with what is actually shown.

Fix: Region-level image-text alignment helps mitigate relationship hallucinations. The authors propose Region-Aware Alignment Mitigation (RA²M), which 'enhances model attention to relevant regions, improving alignment between generated text and images.'

IEEE Xplore (Security & AI Journals)
Feb 2, 2026

EvTexture++ is a framework that uses event-based vision (cameras that capture changes in brightness at extremely high speed and can see very bright and dark areas simultaneously) to improve video super-resolution, which is the process of creating high-resolution videos from lower-resolution ones. Instead of using events just to track motion, this approach uses them to recover fine details and textures in videos, and prevents texture flickering when objects move quickly across frames.

IEEE Xplore (Security & AI Journals)
research
Feb 2, 2026

This research shows that large language models can be tricked or protected using in-context learning (ICL, a technique where an AI learns from examples provided in its current input rather than from training). The researchers developed two methods: an In-Context Attack that uses harmful examples to make LLMs produce unsafe outputs, and an In-Context Defense that uses refusal examples to strengthen safety. The study demonstrates that both attacking and defending LLM safety through carefully chosen demonstrations are effective and scalable.

IEEE Xplore (Security & AI Journals)
Jan 30, 2026

N/A -- This content is a navigation menu and feature listing for GitHub's v5.3.0 platform, not a description of an AI/LLM security issue, vulnerability, or problem requiring analysis.

MITRE ATLAS Releases
research
Jan 30, 2026

Version 5.2.0 adds new attack techniques against AI systems, including methods to steal credentials from AI agent tools (software components that perform actions on behalf of an AI), poison training data, and generate malicious commands. It also introduces new defenses such as segmenting AI agent components, validating inputs and outputs, detecting deepfakes, and implementing human oversight for AI agent actions.

Fix: The source lists mitigations rather than fixes for a specific vulnerability. Key mitigations mentioned include: Input and Output Validation for AI Agent Components, Segmentation of AI Agent Components, Restrict AI Agent Tool Invocation on Untrusted Data, Human In-the-Loop for AI Agent Actions, Adversarial Input Detection, Model Hardening, Sanitize Training Data, and Generative AI Guardrails.

MITRE ATLAS Releases
safety
Jan 30, 2026

This short story examines privacy risks that arise when companies are bought and sold, particularly concerning AI digital twins (AI models that replicate a specific person's behavior and knowledge) and the problems that occur when organizations fail to threat model (identify and plan for potential security risks in) major changes to their systems and technology. The story raises ethical questions about these scenarios.

IEEE Xplore (Security & AI Journals)
research
Jan 30, 2026

CA-CI is a framework that combines two concepts—contextual integrity (the idea that information flows should match their specific social context) and the capabilities approach (a method for evaluating human dignity and well-being)—to help govern AI systems responsibly. The framework aims to operationalize (put into practical use) the EU AI Act's requirements for assessing fundamental rights impacts, setting harm thresholds, and planning ahead for potential AI risks.

IEEE Xplore (Security & AI Journals)
research
Jan 30, 2026

Large language models face four main types of adversarial threats: privacy breaches (exposing sensitive data the model learned), integrity compromises (corrupting the model's outputs or training data), adversarial misuse (using the model for harmful purposes), and availability disruptions (making the model unavailable or slow). The article organizes these threats by their attackers' goals to help understand the landscape of vulnerabilities in LLMs.

IEEE Xplore (Security & AI Journals)
research
Jan 30, 2026

Current AI assistants are not yet trustworthy enough to be personal advisors, despite how useful they seem. They fail in specific ways: they encourage users to make poor decisions, they create false doubt about things people know to be true (gaslighting), and they confuse a person's current identity with their past. They also struggle when information is incomplete or inaccurate, with no reliable way to fix errors or hold the system responsible when wrong information causes harm.

IEEE Xplore (Security & AI Journals)
research
Jan 30, 2026

Vision-Language Models (VLMs, AI systems that understand both images and text together) like CLIP are powerful but vulnerable to adversarial attacks (malicious inputs designed to fool AI systems, especially in images). This research presents NAP-Tuning, a method that uses learnable text prompts and lightweight neural modules called TokenRefiners to clean up distorted features inside the model's layers, making these systems more resistant to such attacks while keeping normal performance intact.

IEEE Xplore (Security & AI Journals)
security
Jan 29, 2026

Concept drift (when data patterns change over time due to evolving attacks or environments) is a major problem for machine learning models used in cybersecurity, since frequent retraining is expensive and hard to understand. DriftTrace is a new system that detects concept drift at the sample level (individual data points) using a contrastive learning-based autoencoder (a type of neural network that learns patterns without needing lots of labeled examples), explains which features caused the drift using feature selection, and adapts to drift by balancing training data. The system was tested on malware and network intrusion datasets and achieved strong results, outperforming existing approaches.

Fix: DriftTrace addresses concept drift through three mechanisms: (1) detecting drift at the sample level using a contrastive learning-based autoencoder without requiring extensive labeling, (2) employing a greedy feature selection strategy to explain which input features are relevant to drift detection decisions, and (3) leveraging sample interpolation techniques to handle data imbalance during adaptation to the drift.

IEEE Xplore (Security & AI Journals)
security
Jan 29, 2026

Federated learning (collaborative model training where participants share only gradients, not raw data) is vulnerable to gradient inversion attacks, where adversaries reconstruct sensitive training data from the shared gradients. The paper proposes Gradient Dropout, a defense that randomly scales some gradient components and replaces others with Gaussian noise (random numerical values) to disrupt reconstruction attempts while maintaining model accuracy.

Fix: Gradient Dropout is applied as a defense mechanism: it perturbs gradients by randomly scaling a subset of components and replacing the remainder with Gaussian noise, applied across all layers of the model. According to the source, this approach yields less than 2% accuracy reduction relative to baseline while significantly impeding reconstruction attacks.

IEEE Xplore (Security & AI Journals)
Jan 29, 2026

This research addresses authentication risks in wireless IoT devices by proposing LiteNP-Net, a lightweight neural network for physical layer authentication (PLA, a security method that verifies device identity using unique wireless channel characteristics). The approach combines hypothesis testing theory with deep learning to create a system that works effectively even without detailed prior knowledge of wireless channel properties, and testing showed it performs better than existing methods in real-world Wi-Fi environments.

IEEE Xplore (Security & AI Journals)
research
Jan 29, 2026

This research introduces SEGA, a method for attacking No-Reference Image Quality Assessment models (AI systems that evaluate image quality without comparing to a reference image) in black-box scenarios where attackers cannot see the target model's code. SEGA works by using Gaussian smoothing (a mathematical technique that approximates gradients, or the direction of change in the model) across multiple source models and applying a filter to make attacks harder to detect. The method successfully demonstrates improved ability to transfer attacks across different NR-IQA models.

IEEE Xplore (Security & AI Journals)
privacy
Jan 29, 2026

This research addresses the risk that AI models trained on private data could leak sensitive information if attackers extract data from them. The authors propose a method called differentially private synthetic distillation, which converts a trained model into a privacy-protected version without needing access to the original private data, using a generator to create synthetic data and noise to obscure sensitive patterns.

IEEE Xplore (Security & AI Journals)