Research
Academic papers, new techniques, benchmarks, and theoretical findings in AI/LLM security.
Academic papers, new techniques, benchmarks, and theoretical findings in AI/LLM security.
16 items
The paper proposes DSW, a dataset-specific watermarking method for detecting excessive use of protected datasets when text-to-image diffusion models are fine-tuned with LoRA. An encoder embeds a watermark image representing the dataset owner's identity into the dataset samples, and a decoder trained with a bi-level optimization strategy extracts that watermark from images generated by the fine-tuned model. The authors report experiments validating the availability, effectiveness, and stealthiness of DSW.
Researchers address long-term spatio-temporal forecasting (LSTF) for applications such as parking resource prediction and environmental quality monitoring. They identify limits of multi-GNNs in this setting, including limited generality, under-utilization of context, static graph merging and overlooked dynamic interrelations. They propose graph structures that encode each node's contextual information, plus a dynamic multigraph fusion architecture with trainable weight tensors, and report significant performance gains for existing GNNs on three large-scale benchmark datasets.
StrokePIN is a PIN authentication system for mobile devices that combines keystroke dynamics with a Siamese Network, a few-shot learning technique that deploys without retraining. The authors collected two new multi-modality datasets covering 20 PINs from 116 users, which they made publicly available. On unseen users the system reports an FAR of 2.2% and an FRR of 1.9%, and entropy analysis puts keystroke dynamics at 4.03-5.83 bits of security against 3 to 10 online guessing attacks.
The authors propose a cross-image federated learning approach for hyperspectral image classification, arguing that single-image processing limits generalization across spatial and temporal domains. The method combines a client-oriented self-guided knowledge-enhanced personalized learning step with a multiscale semantic aligned dynamic aggregation step to reduce bias from uneven data distribution. The authors say it is the first to explore joint HSI classification with federated learning and report results on open-set and closed-set datasets they constructed.
This paper proposes SPARTA, a protocol for avatar authentication in the metaverse that keeps avatars unlinkable to their users. It uses mercurial signatures so users can create multiple unlinkable avatars without repeated registration, and a time-based hash chain so only avatars holding a reputation token can submit feedback on a smart contract. The authors implement it with socket programming on two Raspberry Pi devices, reporting 105 ms end-to-end latency and 17.338 seconds to complete 1000 concurrent authentications (about 57.7 per second).
Ano2Rule is a rule-based method that makes unsupervised anomaly detection models interpretable for security use. It decomposes normal data into compositional distributions using an unsupervised Interior Clustering Tree, then applies Compositional Boundary Exploration to infer the original model's decision boundary on each part. Experiments on network intrusion detection and IoT security datasets reported high fidelity to the original model's predictions.
Fractional-Order Federated Learning presents FOFedAvg, a FedAvg variation that uses fractional-order stochastic gradient descent (FOSGD) to capture long-range relationships and historical information in federated training. The authors compare it against established federated optimization algorithms on nine benchmark datasets across several non-IID partitioning schemes, reporting that it is competitive with and often outperforms these baselines in test performance and convergence speed. They also prove convergence to a stationary point under standard smoothness and bounded-variance assumptions for fractional order 0 < alpha <= 1.
Researchers propose a privacy-enhancing key generation protocol for Attribute-Based Encryption, guided by the Minimal Disclosure principle, so users reveal only the attributes needed for authorization. The protocol decouples attribute verification from key issuance, using reusable verifiable attribute tokens and blinded key requests, with a batch verification mechanism to reduce overhead. The authors report proofs of binding and hiding properties, formal verification in the symbolic model with ProVerif, and experiments showing reduced inference leakage and efficiency gains over existing schemes.
Researchers propose CoDefend, a split learning framework for edge computing that combines local epoch regulation with time-aware detection. The method assigns heterogeneous clients appropriate local epoch numbers and uses a detection window to identify malicious manipulation. Experiments on simulated and NVIDIA Jetson edge platforms show roughly 2× faster training than baselines, with comparable accuracy and effective detection even under collusion.
The paper examines gradient leakage attacks in federated learning, which reconstruct sensitive training data from transmitted gradients. The authors find that privacy information is unevenly distributed across neural network layers, so uniform gradient perturbation is suboptimal. They propose Layer-Specific Gradient Protection (LSGP), which scales protection intensity per layer, and report that it achieves stronger defense than uniform baselines with comparable model utility.
Cert-SSBD is a certified backdoor defense for deep neural networks that replaces fixed randomized smoothing noise with sample-specific noise magnitudes, optimized per sample via stochastic gradient ascent. The method retrains several smoothed models on multiple poisoned training sets and aggregates their predictions, and it introduces a storage-update-based certification method to handle the varying noise. The authors report effectiveness on multiple benchmark datasets and provide code at https://github.com/NcepuQiaoTing/Cert-SSBD.
Researchers propose Rap-LI, a risk-aware privacy preservation framework for LLM inference services such as ChatGPT. It identifies and labels risky content in user prompts, then applies a risk-aware local differential privacy mechanism to sanitize the text. The authors report an average 51.68% improvement in protection against sensitive information leakage over methods with comparable utility.
FAOSPA is a frequency-domain adversarial attack that projects spatiotemporal features from the UCF-Crime dataset into orthogonal subspaces and perturbs frequency bands critical for anomaly discrimination in STEAD, a video anomaly detection model with a baseline ROC-AUC of 0.91. Under minimal perturbation budgets, the attack reduces ROC-AUC to 0.338 and F1-score from 0.824 to 0, achieving a 99% attack success rate that outperforms competing methods by 17.9% with 32.5% lower L2 distortion. The authors argue these results show the threat posed by frequency-domain attacks and the need for more robust defenses.
Researchers propose SIMix, a training framework for multiuser semantic communication that combines Over-the-Air Mixup (OAM) with label-aware user grouping. The OAM mixes users' semantic features over wireless channels to obfuscate sensitive data and cut communication overhead, while an extended max-clique algorithm partitions users to reduce model inversion attack success. On CIFAR-10 and Tiny ImageNet, the method reduces communication overhead by up to 25%, reaches 17.58 dB PSNR under inversion attack (a 20.98 dB reduction), and lowers label inference attack success by 13.44%.
PPOM-Attack is a black-box attack against face recognition that avoids substitute models by taking feedback directly from the target model. It uses a proximal policy optimization (PPO)-based agent to predict and disturb perturbation regions in a face image, and a minimum brightness offsets method to keep adversarial images high quality. The authors report a 21.7% average gain in attack success rate over state-of-the-art FR attacks across seven FR models.
PromptFuzz is a testing framework that applies software fuzzing techniques to assess how robust LLMs are against prompt injection attacks. It runs in two stages, a prepare phase that selects seed prompts and collects few-shot examples, and a focus phase that generates diverse prompt injections. In a real-world competition it ranked 7th of over 4000 participants within 2 hours, and 92% of 50 popular LLM-integrated applications were exploitable with its prompts.