Research
Academic papers, new techniques, benchmarks, and theoretical findings in AI/LLM security.
Academic papers, new techniques, benchmarks, and theoretical findings in AI/LLM security.
22 items
This paper analyzes Local Chained Attacks (LCAs), which span local dataset pollution, selective inputs, and training process tampering in federated learning clients. It proposes AMBER, a three-layer verification framework that uses vector commitments for dataset integrity, a local consistency check for selective input attacks, and TEE-based verification of input-output correlation. Experiments across multiple datasets, models, and attack scenarios report better defense effectiveness than existing approaches with low overhead, including in Non-IID settings.
HookSteg is a lossless image steganography method that uses process hooking to bypass lossy attacks, such as compression and scaling, on social platforms. It identifies affected nodes through tracking and captures lossy attack instructions using syntax-semantic joint features and dynamic taint tracking. The authors report an average 36.63% improvement in extraction accuracy and a 94.34% gain in locating attack instructions versus manual reverse engineering, tested on Youku and Kutumb.
Researchers propose Causal Bootstrapped Alignment (CBA), an unsupervised framework for video-based visible-infrared person re-identification that learns identity representations from unlabeled video tracklets. CBA combines Causal Intervention Warm-up, which suppresses modality- and motion-induced spurious correlations, with Prototype-Guided Uncertainty Refinement, which resolves cross-modality clustering granularity mismatch. Experiments on the HITSZ-VCM and BUPTCampus benchmarks show CBA significantly outperforms existing USL-VI-ReID methods in this setting.
The authors propose PRBPR, a privacy-preserving redactable blockchain that supports policy hiding and revocation for secure data sharing. It uses a hierarchical blinding factor to give chameleon hash trapdoors dynamic expiration, enabling batch revocation and resilience against trapdoor leakage, and extends CP-ABE with lightweight inner-product encoding to conceal access policies. The reported experiments show up to 7.7× higher efficiency and a 66.8% reduction in computational cost over existing redactable blockchain approaches.
DISARM is a joint hardware-software methodology for mitigating runtime side-channel vulnerabilities, which lets attackers extract sensitive data such as encryption keys by measuring execution-time variations. Unlike prior techniques such as PENDULUM and DifFuzzaR, it uses timing values measured on real embedded devices to generate targeted software fixes for C, C++ and Java source code. Validated on 22 standard benchmarks across five embedded and edge devices, it reports lower execution time overhead, code size overhead and better correctness than those tools.
Researchers introduce Updatable Multi-Party Private Set Intersection (UMPSI), a protocol for collaborative threat intelligence where indicators of compromise change constantly. It composes Distributed Point Functions and zero-sharing for lightweight updates, and uses an Oblivious Pseudorandom Function to protect updated items. The authors prove security in the semi-honest model with non-colluding leaders, and report updates up to 377× faster in a LAN and 44.4× faster in a WAN than state-of-the-art MPSI protocols.
LoRASculpt+ is a framework for adapting Multimodal Large Language Models to downstream tasks while preserving general knowledge. It uses sparse LoRA updates, an asymmetric adaptation strategy that assigns different sparsity and capacity to the LLM and connector modules, and regularization terms that steer updates away from critical pretrained regions. The authors report that the method improves generalization and downstream performance across tasks, model scales and backbones, even at high sparsity, and mitigates catastrophic forgetting.
This paper reformulates rotation estimation as a linear model fitting problem without dropping constraints or introducing singularities. It represents rotation motion as a great circle on a quaternion sphere surface and proposes a voting-based method that runs in parallel on GPUs. The method reportedly solves problems with 10^6 data points and a 99% outlier ratio in under 0.5 seconds.
This ACM Computing Surveys article, published in Volume 58, Issue 13 (pages 1-36) in October 2026, addresses hallucinations in large visual and language models. The source text provided contains only the title and publication details, so the research question, method and findings cannot be reported from it.
Researchers present an efficient threshold multi-party private set intersection (T-MPSI) protocol for privacy-preserving user tracking across distributed IoT location datasets. Their approach combines a batch replicated secret sharing private membership test with a one-round secure aggregation algorithm, and is built for the designated k-collusion model. In LAN experiments, the protocol achieves at least a 6x speedup for secure query and a 3x speedup for secure comparison over the state-of-the-art protocol.
This research proposes a quantum-safe single-shot protocol for machine-to-machine authentication and authorization. The protocol is built around lattice-based digital signature algorithms (DSAs) and key encapsulation mechanisms (KEMs), and it runs over insecure channels to establish tokens with forward secrecy. The authors present it as a foundation for identity infrastructure that remains secure after scalable quantum computers exist.
Amber (Adaptive Affinity Memorization with Layer Mutation) is a continual learning method for Multimodal Deepfake Continual Detection (MDCD), the task of detecting new multimodal deepfake techniques as they emerge. The authors argue that vanilla continual learning fails here because artifact feature drift, caused by the cross-modal gap, makes preserved memory unrepresentative and drives forgetting of artifact features. Amber combines Memorization Affinity Estimation, which keeps high- and low-affinity memory, with Knowledge Layer-wise Mutation, which broadens historical memory, and is evaluated with a new metric, initial comprehension capability (ICC). The authors report that extensive experiments show Amber adapts to new deepfake techniques while retaining prior knowledge.
This paper presents the first systematic study of data poisoning attacks in Multimodal Online Federated Learning (MMO-FL), a setting where IoT edge devices train models in a decentralized, real-time way across multiple modalities. The authors give a theoretical analysis of how such attacks degrade learning performance, then propose a detection and mitigation algorithm for MMO-FL systems. Experiments on the UCI-HAR and USC-HAD multimodal datasets show the approach detects and mitigates the attacks.
Fix: The source proposes a novel detection and mitigation algorithm tailored to MMO-FL systems, but does not state a specific fix, patch, configuration change or workaround.
IEEE Xplore (Security & AI Journals)This paper proposes two federated learning frameworks, DRAG and BR-DRAG, to counter client drift from data heterogeneity and Byzantine attacks. DRAG uses a reference direction and a divergence-of-degree metric to linearly calibrate each worker's local update without extra communication. BR-DRAG extends this under Byzantine attacks by keeping a vetted root dataset at the server to produce trusted reference directions, and the authors report analytical convergence proofs for non-convex models plus experiments showing superior performance over state-of-the-art methods.
Ellipsoid Control is a test-time jailbreak defense for large language models that takes a white-list approach instead of relying on collected harmful samples. It runs projected gradient descent to elicit refusal on arbitrary inputs, while an anisotropic ellipsoid fitted from abundant benign data constrains the update to limit distortion of the benign latent geometry. The authors report that across multiple LLMs, jailbreak attacks, benign tasks and safety-boundary evaluations, it improves safety while better preserving utility.
The paper proposes the first differentially private (DP) algorithm for general weighted empirical risk minimization (wERM), where individual contributions to the objective vary, and derives its privacy guarantees and empirical and population utility bounds. It applies this framework to outcome-weighted learning (OWL) for individualized treatment rules. Experiments on simulated and real data show that OWL trained via DP-wERM keeps strong DP guarantees with robust performance.
This survey reviews harmful fine-tuning attacks against large language models and the defenses proposed to counter them. The source text provided is only a citation line (ACM Computing Surveys, Volume 58, Issue 13, pages 1-39, October 2026) and does not describe the survey's method or findings.
This survey reviews metrics for evaluating privacy-preserving generative models. The source text provided is limited to a citation line: ACM Computing Surveys, Volume 58, Issue 13, Pages 1-35, October 2026, with no abstract or article content.
This paper proposes a classification task-oriented method for differentially private data publishing. The method keeps class labels of data records unchanged while achieving differential privacy, uses a large language model to evaluate correlations among category attributes and merges highly correlated ones, and applies clustering to reduce the sensitivity of numerical attributes. Across the Adult, Bank, Statlog and Heart Disease datasets, the method's average F1 score improvements over other methods are 13.34%, 20.30%, 22.94% and 16.55%, respectively.