Research
Academic papers, new techniques, benchmarks, and theoretical findings in AI/LLM security.
Academic papers, new techniques, benchmarks, and theoretical findings in AI/LLM security.
22 items
The source provides only bibliographic metadata for a paper titled "SALT: Semantic-guided adaptive latent space truncation sampling watermarking for diffusion models." It lists the publication date as September 2026, the Journal of Information Security and Applications (Volume 101), and the authors Yuning Guo, Xuewen He, Yue He and Yong Chen. No research question, method, or findings are included.
This paper proposes a unified framework for private set operations (PSO) that removes the intersection-cardinality leakages identified by Guo et al. (USENIX Security 2022) and the during-execution leakage in PSU identified by Jia et al. (USENIX Security 2024). It is built on a new central building block, permuted batched private membership test sharing, with instantiations for balanced and unbalanced settings. Experiments report 0% leakage for PSI-Sum/SS at about 3.3x higher communication for n=2^20 (balanced) and under 2x overhead (unbalanced), while the enhanced PSU achieves up to 1.6x/60.5x lower communication and up to 1.4x/15.6x faster runtime.
The paper introduces TEASE, an algorithm that selects which challenge-response pairs (CRPs) of a PUF to release so that leaking them reveals little about the underlying response model. Selection uses bilevel programming, and experiments on Arbiter, XOR, XRBR, XRRO and Interpose PUFs report ML attack accuracy near 50% on non-leaked CRPs, even with up to 30K CRPs leaked.
Researchers propose GCLC, a Graph-based Contrastive Learning and Clustering framework for classifying encrypted network traffic in open-world settings. Encrypted flows are modeled as traffic interaction graphs and encoded by a Graph Neural Network trained with a hybrid loss, while a Mahalanobis distance-based detector finds unseen classes. Evaluated on 3 public datasets and 1 self-collected dataset, GCLC outperforms existing methods on 4 metrics, with 95.02% open-set new class recognition accuracy.
This paper presents NatIMG-FL, a federated learning framework for classifying X-ray images across healthcare institutions without sharing sensitive data. It uses natural images as auxiliary supervised data to align feature distributions between natural and X-ray images, and introduces a dual weights-based knowledge transfer method to handle heterogeneous model architectures.
The paper proposes Feature-space Planes Searcher (FPS), an unsupervised domain adaptation framework that corrects decision-boundary misalignment while keeping a pre-trained feature encoder frozen. It argues this avoids the cost and unpredictable distortions of fine-tuning, and adds an Intra-Class Distance Metric (ICDM) for label-free hyperparameter selection. The authors report competitive benchmark results and applications including protein structure prediction, remote sensing classification and earthquake detection.
Researchers propose a dataset watermarking framework to detect unauthorized use and trace leaks of image datasets used to fine-tune text-to-image diffusion models such as Stable Diffusion. The framework is assessed on effectiveness, harmlessness, traceability, stealthiness and robustness, including cross-dataset transferability, robustness to image transformations and adaptive attacks, and scalability. Experiments on Stable Diffusion report reliable detection and accurate tracing with minimal data modification while preserving output quality.
The paper proposes Noise-Optimized Adversarial Examples (NOAE), an attack method against unsupervised deep-learning multivariate time series anomaly detection models used in the Industrial Internet of Things. Experiments show NOAE significantly reduces detection precision on industrial benchmarks, including a 71.51% precision drop on the SMAP dataset at a 0.01 noise amplitude. The authors also propose a Hybrid Adversarial Defense (HAD) training approach that uses data-end random segment replacement augmentation to improve robustness.
Fix: The authors propose Hybrid Adversarial Defense (HAD), a training approach that uses adversarial examples and data-end random segments replacement augmentation to alleviate the weak robustness of anomaly detection models.
IEEE Xplore (Security & AI Journals)OpenAI researchers report that reinforcement learning on realistic scenarios targeting beneficial traits, such as honesty, epistemic humility, corrigibility and concern for human welfare, improves alignment across dozens of public and internal evaluations. The gains extend to domains and grading setups absent from training, even when training is restricted to a single domain such as health. The authors also report that the resulting models are harder to steer toward harmful behavior with adversarial prompts or fine-tuning.
TabHGIF is a hypergraph-aware unlearning framework for tabular data that represents rows and attributes as a row-attribute hypergraph to capture high-order feature correlations. It introduces a Hypergraph Influence Function (HGIF) to approximate parameter shifts from structural edits, enabling one-shot parameter correction without re-accessing the raw forget samples. On three benchmarks and four hypergraph backbones, it reaches post-unlearning accuracy close to full retraining, with 2.18–4.27× speedups for row unlearning and 5.05–7.67× for column unlearning, and MIA-AUC values of 0.5014–0.5464.
This paper proposes PQSC, a post-quantum secure semantic communication framework that integrates learning with errors (LWE) encryption into a VQ-VAE-based SemCom system. The authors report that PQSC resists quantum attacks and chosen-plaintext attacks, and that it consistently outperforms baseline methods across datasets, channel conditions and SNR levels. Channel coding and modulation are simulated with Nvidia Sionna.
Double-Blind Cleanser (DBC) is a backdoor unlearning framework for deep neural networks that needs neither a trusted clean dataset nor trigger recovery or poisoned-sample identification. It first uses catastrophic forgetting to remove backdoor behavior, then applies Sharpness-Aware Minimization (SAM) with importance sampling to restore legitimate functionality. The authors report that DBC neutralizes diverse backdoor attacks while maintaining model utility and placed first in the mitigation round of the TrojAI competition.
MicroPatch is a backdoor erasing approach for deep neural networks that targets the model parameters influenced by backdoor triggers. It reconstructs trigger patterns through reverse engineering, uses influence functions to separate victim parameter components from clean ones, and patches those components. The authors report that across four datasets and four representative backdoor attacks, plus spatial-frequency and frequency-domain attacks, MicroPatch reduces attack success rates more than existing methods while keeping high classification accuracy.
Fix: The source describes MicroPatch itself as the proposed mitigation: patching decoupled victim parameter components to purify the model. No separate fix, patched version or configuration change is stated.
IEEE Xplore (Security & AI Journals)This paper formalizes external data extraction attacks (EDEAs) against retrieval-augmented LLMs (RA-LLMs), where sensitive or copyrighted knowledge-base data can be extracted verbatim. The authors propose a framework built from extraction instruction, jailbreak operator, and retrieval trigger, and implement an attack called Secret. Across 4 models, including 3 commercial LLMs, Secret outperforms prior attacks and succeeds against all 16 tested RAG instances, extracting 35% of the data from RAG powered by Claude 3.7 Sonnet where other attacks yield 0%.
Viewpoint Invariant Adversarial Training (VIAT) treats viewpoint change as an attack and trains vision models to resist it by solving a minimax problem with a Gaussian mixture of adversarial viewpoints generated by GMVFool. The authors contribute the ImageNet-V+ benchmark and ViewRS, a certified viewpoint robustness metric, and report that VIAT significantly improves viewpoint robustness across CNN, ViT and multimodal large language model architectures.
Researchers from Boston Children's Hospital's Manton Center, Harvard University and OpenAI used the OpenAI o3 Deep Research model to reanalyze 376 previously unsolved rare-disease cases. The model produced evidence-linked candidate explanations that specialists reviewed, and physicians established diagnoses in 18 cases, an additional diagnostic yield of 4.8%. The study was published June 18, 2026, in NEJM AI.
Researchers ask whether external groups can evaluate frontier language models by substituting WildChat, a public dataset of about 1 million conversations collected between April 2023 and May 2024, for private production data in their Deployment Simulation technique. They report that WildChat-based predictions of real-world failure rates are surprisingly accurate, typically within roughly 3x error for GPT-5.1, 5.2 and 5.4, despite a 2-3 year data gap. Predictive performance degrades most for more technical and agentic forms of misalignment.
This survey reviews Direct Preference Optimization (DPO), an RL-free alternative to Reinforcement Learning from Human Feedback (RLHF) for aligning LLM policy models with human preferences. It covers theoretical analyses, variants, relevant preference datasets, and applications. The authors categorize recent DPO studies by key research questions and propose future research directions.
OpenAI describes Deployment Simulation, a method that replays privacy-preserving samples of past conversations with a new candidate model before release to preview its behavior. The company says the method improved estimates of undesired behavior rates across multiple GPT-5-series Thinking deployments, surfaced novel misalignment, and reduced models' ability to recognize they are being tested. It can measure behaviors occurring at no less than 1 in 200,000 messages.