aisecwatch.com
DashboardVulnerabilitiesNewsResearchArchiveStatsDatasetFor devs
Subscribe
aisecwatch.com

Real-time AI security monitoring. Tracking AI-related vulnerabilities, safety and security incidents, privacy risks, research developments, and policy changes.

Navigation

VulnerabilitiesNewsResearchDigest ArchiveNewsletter ArchiveSubscribeData SourcesStatisticsDatasetAPIIntegrationsWidgetRSS Feed

Maintained by

Truong (Jack) Luu

Information Systems Researcher

Research

Academic papers, new techniques, benchmarks, and theoretical findings in AI/LLM security.

to
Export CSV
1217 items

UQLM: A Python Package for Uncertainty Quantification in Large Language Models

inforesearchPeer-Reviewed
researchsafety
Dec 31, 2025

Hallucinations (instances where Large Language Models generate false or misleading content) are a safety problem for AI applications. The paper introduces UQLM, a Python package that uses uncertainty quantification (UQ, a statistical technique for measuring how confident a model is in its answer) to detect when an LLM is likely hallucinating by assigning confidence scores between 0 and 1 to responses.

Fix: The source describes UQLM as 'an off-the-shelf solution for UQ-based hallucination detection that can be easily integrated to enhance the reliability of LLM outputs.' No specific implementation steps, code examples, or version details are provided in the source text.

JMLR (Journal of Machine Learning Research)

Nonparametric generative modeling for time series via Schr{\"{o}}dinger bridge

inforesearchPeer-Reviewed
research

On the Relevance of Byzantine Robust Optimization Against Data Poisoning

inforesearchPeer-Reviewed
research

Statistical guarantees for denoising reflected diffusion models

inforesearchPeer-Reviewed
research

Robust training of implicit generative models for multivariate and heavy-tailed distributions with an invariant statistical loss

inforesearchPeer-Reviewed
research

Abstract Gradient Training: A Unified Certification Framework for Data Poisoning, Unlearning, and Differential Privacy

inforesearchPeer-Reviewed
research

From Zipf's Law to Neural Scaling through Heaps' Law and Hilberg's Hypothesis

inforesearchPeer-Reviewed
research

Breaking the Curse of Dimensionality: Diffusion Models Efficiently Learn Low-Dimensional Distributions

inforesearchPeer-Reviewed
research

Unveiling the Statistical Foundations of Chain-of-Thought Prompting Methods

inforesearchPeer-Reviewed
research

AgentPEN: A Prediction-Explanation Network for Sequential Stock Movement via LLMs and Recurrent Generation

inforesearchPeer-Reviewed
research

MarkDiffusion: An Open-Source Toolkit for Generative Watermarking of Latent Diffusion Models

inforesearchPeer-Reviewed
research

Dialing for Dollars or Defaulting Online? Assessing Borrower Risk through Call Activity and Social Media Engagement in Microfinance

inforesearchPeer-Reviewed
research

Adoption of ChatGPT in Organizations: Technology Affordance and Constraints Theory Perspective

inforesearchPeer-Reviewed
research

Enhancing Robustness in Deep Convolutional Neural Networks Through Multiresolution Learning

inforesearchPeer-Reviewed
research

HGNN Shield: Defending Hypergraph Neural Networks Against High-Order Structure Attack

inforesearchPeer-Reviewed
security

Source-Free Time-Series Domain Adaptation With Prior Evaluation of Model Salience

inforesearchPeer-Reviewed
research

Neural Machine Unranking

inforesearchPeer-Reviewed
research

Large Language Models in Human Subject Research, and the Presence of Idiosyncratic Human Behaviors

inforesearchPeer-Reviewed
research

Generative Artificial Intelligence: Ethical Challenges and Trust Mechanisms

inforesearchPeer-Reviewed
research

The Impact of Artificial Intelligence in Protecting the Online Social Community From Cyberbullying

inforesearchPeer-Reviewed
research
Previous51 / 61Next
Dec 31, 2025

Researchers propose a generative model for time series (sequences of data points over time) based on Schrödinger bridge, a mathematical technique that uses optimal transport (finding the most efficient way to transform one data distribution into another) to create synthetic time series data. The model estimates unknown functions from real data using nonparametric methods (techniques that don't assume a specific underlying mathematical form) and generates new synthetic samples that preserve the temporal patterns in the original data.

JMLR (Journal of Machine Learning Research)
security
Dec 31, 2025

This research examines Byzantine robust optimization, a technique that protects machine learning systems when data is poisoned (corrupted or maliciously altered) and some workers (computers processing parts of the dataset) behave unpredictably in distributed networks. The study proves that Byzantine-robust approaches provide optimal protection even when facing weaker threats where only local datasets are poisoned, and shows that having some workers with completely corrupted data is more damaging than having workers with partially corrupted data.

JMLR (Journal of Machine Learning Research)
Dec 31, 2025

This research paper analyzes denoising reflected diffusion models, which are a type of generative AI (systems that create new data like images or text). The study shows that reflected diffusion processes (a mathematical technique using boundaries to keep the model's state space bounded) can match theoretical predictions better than standard diffusion models, and provides mathematical proof of how quickly these models converge to accurate results.

JMLR (Journal of Machine Learning Research)
Dec 31, 2025

This research addresses problems with training implicit generative models (AI systems that learn to create new data similar to real data) by proposing the invariant statistical loss (ISL), which avoids unstable adversarial training by comparing the statistical ranks of real and generated samples. The authors improve ISL for two practical scenarios: using a Pareto distribution instead of Gaussian noise to better model extreme values in data, and introducing ISL-slicing to handle large multivariate datasets (data with many variables) by projecting onto random lower-dimensional subspaces.

JMLR (Journal of Machine Learning Research)
security
Dec 31, 2025

This research introduces Abstract Gradient Training (AGT), a framework for verifying that machine learning models remain reliable when their training data is changed or corrupted. The framework addresses three scenarios: adversarial data poisoning (when attackers intentionally alter training samples), machine unlearning (when specific training data must be removed), and differential privacy (when individual data points are substituted). AGT works by establishing mathematical bounds on how model parameters (the internal settings that AI uses to make predictions) can change, allowing researchers to formally prove a model will behave safely despite training data perturbations.

JMLR (Journal of Machine Learning Research)
Dec 31, 2025

This paper explores the mathematical connection between Zipf's law (a principle stating that word frequencies follow a power law distribution, where a few words appear very often and most appear rarely) and the neural scaling law (which describes how an AI model's prediction error improves as you give it more training data, parameters, or computing power). The authors show that under certain assumptions, the neural scaling law can be derived as a logical consequence of Zipf's law, connecting it through two intermediate principles: Heaps' law (about vocabulary growth) and Hilberg's hypothesis (about information content scaling).

JMLR (Journal of Machine Learning Research)
Dec 31, 2025

This research paper explains how diffusion models (AI systems that generate images by gradually removing noise) can learn data patterns efficiently without being overwhelmed by the curse of dimensionality (a problem where learning becomes exponentially harder as the number of features increases). The researchers show that when data has natural low-dimensional structure (like how real images can be represented with fewer underlying features than their total pixel count), diffusion models can learn distributions with sample complexity that scales linearly with intrinsic dimension rather than exponentially with the space's full size.

JMLR (Journal of Machine Learning Research)
Dec 31, 2025

This research paper analyzes Chain-of-Thought prompting (a technique where AI models show their reasoning steps to solve complex problems) from a statistical perspective, rather than just observing that it works. The authors create a mathematical model showing that CoT prompting approximates a Bayesian estimator (a statistical method for making predictions based on prior knowledge and examples), and they prove that the error in this approach comes from two sources: difficulty in understanding the prompt itself, and limitations in the pretrained language model. The paper demonstrates that providing more examples in the prompt reduces the first type of error exponentially.

JMLR (Journal of Machine Learning Research)
Dec 31, 2025

AgentPEN is a system that uses large language models (LLMs, AI systems trained on massive text data) to predict stock price movements while explaining why those predictions are made. The system works by combining financial news data with stock price information through a special agent that selects relevant news, remembers important information over time and across different sources, and then predicts stock movements based on the combined insights.

JMLR (Journal of Machine Learning Research)
Dec 31, 2025

MarkDiffusion is an open-source Python toolkit that helps researchers add watermarks to latent diffusion models (AI systems that generate images from text descriptions). The toolkit includes tools for adding watermarks, visualizing how they work, and testing whether watermarks remain detectable and effective even if someone tries to remove them.

JMLR (Journal of Machine Learning Research)
Dec 31, 2025

This research studies how to predict whether borrowers on micro-lending platforms (small-loan services) will default (fail to repay their loans) by examining their call activity and social media behavior. The study analyzed over 154,000 loans from Indonesian platforms and found that frequent calls and stable calling patterns suggest lower default risk, while frequent social media activity and stable social media patterns actually indicate higher default risk. These findings suggest that micro-lending platforms could improve their credit assessment models (systems for deciding who gets loans) by combining both types of behavioral data.

AIS eLibrary (Journal of AIS, CAIS, etc.)
Dec 31, 2025

This research studied what makes knowledge workers (people whose jobs involve handling information) want to use ChatGPT at work, using technology affordance and constraints theory (a framework explaining how tools enable certain actions while limiting others). The study found that ChatGPT's benefits like automation, information quality, and productivity boost adoption, but concerns about risk and lack of regulation reduce it. Personal innovativeness (how open someone is to new ideas) and supportive workplace culture help workers embrace ChatGPT despite their concerns.

AIS eLibrary (Journal of AIS, CAIS, etc.)
safety
Dec 29, 2025

This research explores multiresolution learning, a training method where AI models learn from data at multiple levels of detail, starting from very coarse versions and progressively moving to finer ones. The study shows this approach makes deep neural networks (DNNs, which are AI systems with many layers) more robust against noise and adversarial attacks (deliberate attempts to fool the AI) while maintaining accuracy, without requiring extra computing power compared to traditional training methods.

IEEE Xplore (Security & AI Journals)
research
Dec 26, 2025

Hypergraph Neural Networks (HGNNs, which are AI models that learn from data where connections can link multiple items together instead of just pairs) can be weakened by structural attacks that corrupt their connections and reduce accuracy. HGNN Shield is a defense framework with two main components: Hyperedge-Dependent Estimation (which assesses how important each connection is within the network) and High-Order Shield (which detects and removes harmful connections before the AI processes data). Experiments show the framework improves performance by an average of 9.33% compared to existing defenses.

Fix: The HGNN Shield defense framework addresses the vulnerability through two modules: (1) Hyperedge-Dependent Estimation (HDE) that 'prioritizes vertex dependencies within hyperedges and adapts traditional connectivity measures to hypergraphs, facilitating precise structural modifications,' and (2) High-Order Shield (HOS) positioned before convolutional layers, which 'consists of three submodules: Hyperpath Cut, Hyperpath Link, and Hyperpath Refine' that 'collectively detect, disconnect, and refine adversarial connections, ensuring robust message propagation.'

IEEE Xplore (Security & AI Journals)
Dec 24, 2025

This paper addresses source-free domain adaptation (SFDA, a technique that adapts AI models to new datasets without accessing the original training data) for time-series data, such as sensor readings or activity logs. The authors argue that existing methods lack interpretability and may learn spurious patterns, so they propose PrEPoA, a framework that evaluates which parts of the time-series data the model considers important before fine-tuning it on the target domain. They demonstrate their approach works better than existing methods across five different real-world datasets.

IEEE Xplore (Security & AI Journals)
privacy
Dec 23, 2025

This research addresses machine unlearning in neural IR (information retrieval, the technology that ranks search results), a process called neural machine unranking (NuMuR) that selectively removes data from AI systems for privacy compliance. The authors propose CoCoL (contrastive and consistent loss, a method with two complementary training objectives), which uses a contrastive loss to reduce relevance scores on forgotten data while preserving performance on shared data, plus a consistent loss to maintain accuracy on retained data, demonstrating effective data removal across multiple neural ranking models.

Fix: The proposed solution is CoCoL, a dual-objective framework comprising: 1) a contrastive loss that reduces relevance scores on forget sets while maintaining performance on entangled samples, and 2) a consistent loss that preserves accuracy on the retain set. According to the paper, CoCoL achieves substantial forgetting with minimal retention and generalization performance loss.

IEEE Xplore (Security & AI Journals)
safety
Dec 22, 2025

Large language models (LLMs, AI systems trained on huge amounts of text to generate human-like responses) can now mimic not just general human language but also unusual, individual-specific human behaviors. This ability could lead to LLMs being used more widely in research studies and potentially reduce the role of actual humans, which raises concerns about AI alignment (ensuring AI systems behave in ways humans intend and approve of) and how this technology affects society.

IEEE Xplore (Security & AI Journals)
safety
Dec 22, 2025

Generative AI (systems that create new text, images, or other content) is transforming many industries but raises ethical concerns like data privacy (protecting personal information), bias (unfair treatment of certain groups), transparency (being open about how the AI works), and accountability (responsibility for the AI's actions). Researchers propose a trust framework based on transparency, fairness, accountability, and privacy to help ensure generative AI is developed and used responsibly.

IEEE Xplore (Security & AI Journals)
safety
Dec 22, 2025

Cyberbullying on social media is a growing problem that harms people's mental health, and traditional methods to stop it are no longer effective. This study examines how artificial intelligence can help protect online communities from cyberbullying by exploring different AI technologies, their uses, and the challenges involved. The goal is to understand how AI might create safer online environments.

IEEE Xplore (Security & AI Journals)