Research
Academic papers, new techniques, benchmarks, and theoretical findings in AI/LLM security.
Academic papers, new techniques, benchmarks, and theoretical findings in AI/LLM security.
45 items
This paper presents a reinforcement learning framework that guarantees robustness under bounded uncertainties in dynamics modeling and state estimates. The uncertainties and system dynamics are described with ReLU neural networks, which let the method find the states that most violate the robustness conditions and prioritize them during training. The authors evaluate the approach in simulation, including 3-D quadrotor reference trajectory tracking.
Learning Optimal Policies With Local Observations for Cooperative Multiagent Reinforcement Learning proposes unified MARL (UMARL), a weighted value function factorization method that combines exploitation and exploration in one framework. The authors prove a latent state can guarantee optimal individual and global policies and can be approximately obtained from local observations. UMARL uses an agent representation network, individual weighting networks and a latent state regularizer, and outperforms 12 state-of-the-art methods on m-step matrix game, level-based foraging, StarCraft II and Google research football.
Researchers propose preferred-label partial-label learning in a one-vs-all view (PLL-OVA), where each training instance has a candidate label set plus a preferred label within it. They model preferred-label generation with a noisy-channel formulation covering symmetric and asymmetric mis-selection and missing true labels, and build empirical risk minimization procedures with risk-correction functions such as ReLU and ABS. Experiments on benchmark datasets report that PLL-OVA consistently outperforms standard PLL baselines, especially under large candidate ambiguity or nonuniform mis-selection.
Researchers Zhongyuan Zhang, Xinrui Ge and Jia Yu published a paper in Computers & Security, available online 18 March 2026, on privacy-preserving spatio-temporal keyword queries with verifiability for location-based services. The source text provided contains only the publication metadata, so the research question, method and findings cannot be described from it.
Researchers introduce sparse variational Student-t processes (SVTPs), a sparse inducing point framework that extends Student-t processes to large datasets, with two inference algorithms, SVTP-UB and SVTP-MC. On UCI and Kaggle datasets with outliers and heavy tails, SVTP reports up to 3× faster convergence and 40% lower prediction error than sparse GPs, while remaining tractable for datasets with over 200 000 samples.
This survey reviews learning-based dynamic fault localization, covering methods from traditional machine learning through to large language models. It appears in ACM Computing Surveys, Volume 58, Issue 9, pages 1-39, July 2026.
Researchers propose AdaptiveGDE, a differential evolution algorithm for high-dimensional nonconvex optimization. It uses a two-step mutation operator that separates differential mutation from gradient descent, and an adaptive niching strategy that adjusts subpopulation count based on population similarity and iteration progress. Under relaxed smoothness assumptions and approximate ℓ1 regularization, the authors prove convergence in expectation to a near-optimal solution within O(1/ε^4) iterations, and report improved test accuracy and loss in DNN training, especially with limited data.
Researchers propose a novel Android malware detection method built on CWInFs and MPTACF optimization. The source text is an introduction that surveys the field, covering syntactic features such as permissions and API calls, DREBIN-style feature weighting, classifier ensembles, and data imbalance handling. It does not yet report the method's results.
Researchers Bruce W. Lee, Yueh-Han Chen and Tomek Korbak train agents to call a report_scheming() tool whenever they covertly misbehave, a method they call self-incrimination. For GPT-4.1, undetected successful attacks fall from 56% to 6%, outperforming matched-capability blackbox monitors and alignment baselines across 15 out-of-distribution environments. The approach also preserves general capabilities and generalizes from instructed to uninstructed misbehavior.
PureDiffusion is a dual-purpose framework for diffusion models that inverts backdoor triggers for defense and reinforces them for attack amplification. Its defense uses two novel loss functions based on trigger-induced distribution shifts across timesteps and denoising consistency, followed by a detection method. Reported results include near-perfect detection accuracy and attack success rates of nearly 100% for existing backdoor attacks, with training time reduced by up to 20x.
This paper proposes a backdoor defense for vertical federated learning (VFL), where an attacker controls its own local data and model and the defender cannot see them. The authors introduce a latent masked autoencoder (LMAE) that measures the semantic consistency of embeddings from different VFL participants, since malicious embeddings tend to be less predictable from honest ones. The defense identifies attackers and enables backdoor-resistant predictions, and the authors report it is effective across four baseline defenses, seven backdoor attacks and five datasets of different modalities.
The OWASP GenAI Security Project, an open-source community it describes as growing to more than 25K members, announced new resources and a full week of RSA programming ahead of RSA 2026, dated March 19, 2026. The announcement also highlights growing industry adoption and continued sponsor support.
The paper proposes FORCE, a Byzantine-resilient aggregation scheme for Decentralized Federated Learning (DFL), where clients share models directly without a central server. Existing defenses rely on gradient checks that adaptive adversaries can manipulate, so FORCE instead uses the loss of each local model, inspired by the Shapley value from game theory, to flag potentially malicious clients. A lightweight variant, FORCE⁻, approximates the Shapley computation to reduce overhead as neighbor counts grow, and experiments on four datasets under three attacks report that FORCE outperforms existing state-of-the-art methods.
Researchers propose a Two-Stage Defense Framework (TSDF) against deepfakes that aims to keep active defenses effective after attackers retrain their models. The framework uses dual-function adversarial perturbations that distort forged results and also poison the data preparation step of an attacker's retraining pipeline. The source reports that traditional interruption methods degrade sharply under adversarial retraining, while TSDF shows stronger dual defense capability.
OpenAI published a research report on how Americans use ChatGPT for wage information, finding nearly 3 million wage-related messages per day in the US. The most common requests were pay calculation (26%), specific role pay (19%), and entrepreneurship (18%). The report also introduces WorkerBench, which evaluated GPT-5.4 against 2024 OEWS median wages and found high coverage, small bias, and numeric estimates very close to the benchmark.
This ACM Computing Surveys article, published July 2026 in Volume 58, Issue 9 (pages 1-37), examines alignment of diffusion models. The source text provided contains only the citation details and no article content, so no methods or findings can be described.