Academic papers, new techniques, benchmarks, and theoretical findings in AI/LLM security.
This paper describes a new method for detecting AI-generated images (images created by GANs, which are machine learning models that generate synthetic images, or diffusion models, which gradually refine noise into images) by analyzing images in multiple frequency domains (different ways of breaking down an image into mathematical components) using attention mechanisms (techniques that help AI focus on important parts of data). The approach achieved better detection accuracy than previous methods when tested on images from 65 different generative models.
Website fingerprinting (WF) attacks are methods used to identify which websites a person visits even when they use Tor encryption (a privacy tool that hides browsing activity). Existing attacks work well when someone visits one website at a time, but struggle when multiple website tabs are open simultaneously. This research presents STMWF, a new attack that combines spatial-temporal sequence analysis (examining the order and timing of data packets sent between a user's computer and websites) with machine learning techniques to better identify websites even when multiple tabs are open, showing significant improvements over previous methods.
This research studies how small and medium-sized companies decide whether to build their own digital platform or join an existing one, using Resource Dependence Theory (a framework explaining how organizations manage their needed resources). The study found that companies worry more about becoming dependent on platforms than about lacking resources, and that data dependence (reliance on information controlled by platforms) is a new and important factor that traditional theories didn't account for.
This research examines why individuals do not widely adopt personal cyber insurance, which covers remaining risks that preventive security measures cannot stop. Using survey data from 301 U.S. residents and analyzing cognitive factors through fsQCA (fuzzy-set qualitative comparative analysis, a method that identifies different combinations of conditions leading to the same outcome), the study finds that different psychological and behavioral factors lead people to either adopt or reject cyber insurance in ways that differ from previous research.
FedPerX is a federated transformer framework (a system where multiple computers train an AI model together without sharing raw data) designed for sentiment analysis across multiple languages while protecting privacy. It uses residual adapters (lightweight customizable modules added to a shared language model) and differential privacy (a mathematical technique that adds noise to data to prevent identifying individuals) to let each participant personalize their model without compromising data privacy. The framework outperforms existing methods on multilingual datasets with improved accuracy and significantly reduced communication needs.
ADVersa is a framework that uses AI to understand and explain traffic accidents by analyzing video and text together. It can recover what happened before a crash, predict what will happen during a crash, and generate explanations for why accidents occur by learning from a new dataset (MM-AU) containing nearly 12,000 accident videos with detailed descriptions and object annotations.
This paper proposes RCMCL (Robust Trusted Conflictive Multiview Collaborative Contrastive Learning), a method to improve AI models that learn from multiple sources of data (multiview learning) when those sources conflict or misalign with each other. The approach uses evidential deep neural networks (a technique that estimates uncertainty in predictions) and contrastive learning (a training method that teaches the model to recognize similar and different examples) to make the model more reliable and accurate even when the data sources provide contradictory information.
Fingerprint authentication systems can be tricked by presentation attacks (using fake fingerprints to gain unauthorized access), and while deep learning methods help detect these spoofs, they struggle because different fake materials look different and the AI focuses on easy cases while missing hard ones. This paper proposes EasyHard-FSD, a method that uses hard sample mining (a technique where an AI identifies the trickiest examples to learn from) with a teacher/student model setup to help the AI better distinguish real fingerprints from fake ones.
AdvScan is a method for detecting adversarial examples (inputs slightly modified to trick AI models into making wrong predictions) on tiny machine learning models running on edge devices (small hardware like microcontrollers) without needing access to the model's internal details. The approach monitors power consumption patterns during the model's operation, since adversarial examples create unusual power signatures that differ from normal inputs, and uses statistical analysis to flag suspicious inputs in real-time with minimal performance overhead.
Researchers have developed a new backdoor attack method called shell code injection (SCI) that can implant malicious logic into deep learning models (neural networks trained on large datasets) without needing to poison the training data. The attack uses techniques inspired by nature, like camouflage, along with trigger verification and code packaging strategies to trick models into making wrong predictions, and it can adapt its attack targets dynamically using large language models (LLMs) to make it more flexible and harder to detect.
This research introduces PP-DR, a privacy-preserving dimensionality reduction (a technique that reduces the number of features in a dataset to make it easier to analyze) scheme that uses homomorphic encryption (a type of encryption that allows computations on encrypted data without decrypting it first) to let multiple organizations securely share and analyze data together without revealing sensitive information. The new method is much faster and more accurate than previous approaches, achieving 30 to 200 times better computational efficiency and 70% less communication overhead.
This paper presents AMD-GCL, a new method for graph contrastive learning (GCL, a technique where an AI learns from graph data by comparing different modified versions of the same graph). Unlike existing approaches that create modifications independently, AMD-GCL uses adversarial augmentation (deliberately adding perturbations to make the modifications more different from each other) to maximize the differences between paired graph modifications, which improves learning performance.
Differentially private databases (DP-DBs, systems that add mathematical noise to data to protect individual privacy while allowing useful analysis) need auditing services to verify they actually protect privacy as promised, but current approaches don't handle database-specific challenges like varying query sensitivities well. This paper introduces DPAudit, a framework that audits DP-DBs by generating realistic test scenarios, estimating privacy loss parameters, and detecting improper noise injection through statistical testing, even when the database's inner workings are hidden.
Fix: The source presents DPAudit as a framework solution but does not describe a patch, update, or deployment fix for existing vulnerable systems. N/A -- no mitigation discussed in source.
IEEE Xplore (Security & AI Journals)PROTheft is a model extraction attack (a method where attackers steal an AI model's functionality by observing its responses to many input queries) that works on real-world vision systems like autonomous vehicles by projecting digital attack samples onto a device's camera. The attack bridges the gap between digital attacks and physical-world scenarios by using a projector to convert digital inputs into physical images, and includes a simulation tool to predict how well attack samples will work when converted from digital to physical to digital formats.
This paper presents KSIQA, a new machine learning model that measures image quality without needing a reference image to compare against (no-reference image quality assessment, or NR-IQA). The model uses a knowledge-sharing strategy where a teacher model that can see reference images helps train a student model to predict quality by generating mental imagery and combining different types of feature extraction (vision transformers, which break images into patches for analysis, and convolutional neural networks, which apply filters to detect patterns). The researchers show their model performs better than existing no-reference quality assessment methods on standard test datasets.
This paper introduces A2Net, a system for whole-body pose estimation (predicting the locations of keypoints on a person's face, body, hands, and feet from an image) that combines vision and language models to solve two problems: scale variation (different body parts appearing at different sizes) and semantic ambiguity in small-scale parts (difficulty identifying what small features represent). The approach uses text features alongside image features because text is not affected by image scaling issues, then aligns them using optimal transport (a mathematical method for matching distributions) to create a unified visual-language representation that improves keypoint localization accuracy.
Version 5.4.0 (released February 5, 2026) is an update to a security framework that documents new attack techniques targeting AI agents, including publishing poisoned AI agent tools (malicious versions of legitimate tools), escaping from AI systems to access the host computer, and exploiting vulnerabilities to steal credentials or evade security. The update also includes new real-world case studies showing how attackers have compromised AI agent control systems and used prompt injection (tricking an AI by hiding commands in its input) to establish control.
This research presents a semisupervised domain adaptation method (SDM), which helps AI classifiers work better when transferring knowledge from one data domain to another (like using a model trained on one type of data to work with a different type). The method addresses two main problems: limited labeled training data in the target domain and distribution divergence (differences in data patterns between source and target domains) by iteratively updating training data while balancing multiple objectives like structural risk and manifold consistency (geometric patterns in data).
This research presents an evidential deep neural network (EDNN), which is a machine learning model that combines evidence theory (a method for handling uncertainty) with convolutional neural networks (CNNs, algorithms that process images). Unlike traditional classifiers that assume all possible categories are known, the EDNN works under an open-world assumption (acknowledging that unknown categories may exist) and can classify items into single categories, multiple possible categories, or identify them as completely novel.
This article presents a method for teaching AI systems to control complex machinery safely and efficiently by combining optimal control (finding the best way to manage a system), safety constraints (hard limits that must never be broken), and a self-triggered mechanism (a system that only updates calculations when necessary). The approach uses a safety-filter structure to balance the competing goals of optimal performance and guaranteed safety, while reducing the computing power needed.