Research
Academic papers, new techniques, benchmarks, and theoretical findings in AI/LLM security.
Academic papers, new techniques, benchmarks, and theoretical findings in AI/LLM security.
21 items
This article presents a reinforcement learning scheme for safe optimal formation tracking of multiple fixed-wing UAVs under external disturbances and asymmetric control constraints. The method builds a safe set from a control barrier function, integrates it into a nonquadratic cost of a constrained zero-sum differential game, and learns the robust safe Nash policy with a critic-only RL strategy using experience replay, avoiding the persistence of excitation condition. The authors verify stability and forward invariance of the safe set and validate the scheme in simulation.
The authors study the sample complexity of stochastic convex optimization when parameters such as the distance to optimality and the Lipschitz constant are unknown. They propose a reliable model selection method that tunes learning rates to match known-parameter sample complexity up to $\log\log$ factors, and a regularization-based method using norm-regularized empirical risk minimization that estimates the distance to optimality within a constant factor. Few-shot experiments on CIFAR-10 with fine-tuned CLIP models and prompt-engineered Gemini shape counting indicate the model selection method can help mitigate overfitting to small validation sets.
The authors propose a nonparametric generative model for time series based on the Schrödinger bridge (SB) approach. The model is characterized by a finite-horizon stochastic differential equation with a path-dependent drift, which the authors estimate from data with kernel regression and simulate to produce synthetic samples. They evaluate it on autoregressive models, a GARCH model and fractional Brownian motion, and apply the synthetic samples to deep hedging on real datasets.
This paper analyzes Chain-of-Thought (CoT) prompting from a statistical perspective, using a multi-step latent variable model to show that the CoT estimator approximates a Bayesian estimator that infers the posterior from examples in the prompt. The authors prove the statistical error splits into a prompting error, which decreases exponentially as the number of prompt examples grows, and a pretraining error of the pretrained LLM. They also construct a transformer model class and establish its generalization error under the PAC-Bayes framework.
This paper derives the neural scaling law, which describes how a foundation model's cross entropy rate changes with training tokens, parameters, and compute, from Zipf's law on token distributions. The authors chain the derivation through Heaps' law on vocabulary growth and Hilberg's hypothesis on entropy scaling. They illustrate the steps with a toy example, the Santa Fe process, which satisfies all four statistical laws.
Adversarially trained implicit generative models can suffer unstable dynamics and mode collapse. This work formally characterizes the invariant statistical loss (ISL) as a proper divergence over continuous distributions, showing it is continuous and differentiable so it supports stable gradient-based optimization without adversarial games. It also introduces Pareto-ISL, which uses a generalized Pareto latent distribution to better model heavy tails, and ISL-slicing, which scales the loss to high-dimensional data by averaging rank-based losses over random one-dimensional projections.
This paper studies statistical guarantees for denoising reflected diffusion models, which use a reflected diffusion process as the noise driver to avoid the thresholding procedures that implementations add to handle the unbounded state space of idealised models. Under Sobolev smoothness assumptions, the authors establish convergence rates in total variation that, up to a polylogarithmic factor, match the minimax lower bound. Their contributions include the statistical analysis of this model class and a refined score approximation method in time and space based on spectral decomposition and neural network analysis.
This paper studies how knowledge transfer affects the generation accuracy of generative models. It fine-tunes a target-task model from a pre-trained source-task model, using a Shared Embedding concept and a transfer learning framework measured by distribution metrics such as Kullback-Leibler divergence. Theoretical and experimental results for diffusion and normalizing flow models show better performance than non-transfer counterparts.
The paper treats diffusion models as an implicit approach to nonparametric density estimation and analyzes them within a statistical framework. It assumes the underlying density factorizes into low-dimensional components, as in Bayesian networks and Markov random fields. The authors show that a diffusion-based implicit estimator adapts to this factorization and achieves the minimax optimal rate under total variation distance, using a sparse weight-sharing neural network architecture.
Meta-ARVDM is a unified analytical framework that studies history forgetting and temporal degradation in Auto-Regressive Video Diffusion Models (AR-VDMs). The authors show that history forgetting is characterized by the conditional mutual information between the generated output and preceding frames, and prove that adding more past frames monotonically alleviates it. They also show that temporal degradation can be quantified by the cumulative sum of per-step errors, and report a strong empirical correlation between the two phenomena.
This paper sets out a mathematical foundation for denoising Markov models, a class of generative models pairing a forward process from the target distribution to a simple distribution with a constructed backward process for sampling. Using connections to nonequilibrium statistical mechanics and generalized Doob's h-transform, the authors give a unified variational objective and a recipe for designing such models driven by arbitrary Lévy-type processes. They demonstrate novel models using geometric Brownian motion and jump processes as forward dynamics.
The paper introduces a two-step method for probabilistic rainfall downscaling, which generates finer-scale rainfall distributions from coarse-scale deterministic weather variables while preserving temporal and spatial dependence. Marginal location-specific distributions are fitted with joint generalised neural models, which extend generalised linear models with a deep neural network, and spatial coherence is modelled with a censored latent Gaussian copula whose correlation matrix comes from a Gaussian Process Kernel. The authors report that the approach outperforms existing methods on a large UK dataset.
The paper develops a mathematical framework explaining how diffusion models learn low-dimensional data distributions from a finite number of training samples without the curse of dimensionality. It models data as a mixture of low-rank Gaussians and shows that the diffusion training objective is equivalent to subspace clustering over the training samples, so sample complexity scales linearly with intrinsic dimension rather than exponentially with ambient dimension. Experiments on synthetic and real image datasets show phase transitions in generalization.
This study examines whether borrowers' call activity and online social activity predict default on micro-lending platforms, drawing on social capital theory's bonding and bridging capital. Using 154,579 loan records from 10 Indonesian micro-lending platforms and multiple robustness tests, the authors find that call activity frequency and stability significantly reduce default risk. Online social activity frequency and stability, by contrast, significantly increase it.
Abstract Gradient Training (AGT) is a unified framework for certifying a given model and training procedure against training data perturbations, covering bounded perturbations, removal of data points, and addition of new samples. It works by bounding the reachable set of parameters to establish provable parameter-space bounds, and it targets models trained with first-order optimization methods. The framework is presented as covering data poisoning, machine unlearning, and differential privacy.
MarkDiffusion is an open-source Python toolkit for generative watermarking of latent diffusion models. It provides a unified implementation framework, a mechanism visualization suite for embedded and extracted watermark patterns, and an evaluation module with 24 tools and 8 automated pipelines for detectability, robustness, and output quality. The counts reflect the initial release, and the code is available at https://github.com/THU-BPM/MarkDiffusion.
AgentPEN is a prediction-explanation network for stock movement that aligns financial news text with price streams using an LLM-based Representation Fusion Agent. It feeds the fused representation into a Deep Recurrent Generation module to predict stock movements. The authors report that it surpasses state-of-the-art baselines in prediction accuracy and explainability on multiple real-world datasets.
UQLM is a Python package for detecting hallucinations in Large Language Models using uncertainty quantification (UQ) techniques. It provides a suite of UQ-based scorers that produce response-level confidence scores from 0 to 1, offered as an off-the-shelf library that can be integrated to improve the reliability of LLM outputs.
The paper examines whether Byzantine machine learning, where distributed workers may deviate arbitrarily from the algorithm, is relevant to data poisoning, a weaker threat model where only local datasets are corrupted. The authors prove that Byzantine-robust schemes yield optimal solutions under both fully-poisonous and partially-poisonous local data, and that fully-poisonous workers are more harmful when local data is heterogeneous.
This study examines what drives knowledge workers to intend to adopt ChatGPT at work. Drawing on technology affordance and constraints theory, focus groups, and a survey, the authors identify three affordances (automatability, information quality, and productivity) and two constraints (perceived risk and lack of regulation). The affordances raise perceived effectiveness, while lack of regulation increases discomfort. Personal innovativeness strengthens the link between effectiveness and adoption intention and weakens the effect of discomfort.