Research
Academic papers, new techniques, benchmarks, and theoretical findings in AI/LLM security.
Academic papers, new techniques, benchmarks, and theoretical findings in AI/LLM security.
25 items
Researchers Saleh Alzahrani, Yang Xiao and Sultan Asiri published RanDS, an open dataset of raw ransomware binaries and extracted features, in Computers & Security, available online 25 March 2026. The source text contains only the title, publication date, journal and authors, so the summary is limited to these bibliographic details.
Researchers propose MoEC, a multiobjective ensemble-critic reinforcement learning method for autonomous driving. Each critic in the ensemble follows an independent reward function, and the method pairs a hybrid parameterized action space with an uncertainty-based exploration mechanism that supports hybrid actions. In simulator-based and HighD dataset-based multilane highway scenarios, the method learns policies that balance efficiency, action consistency, and safety.
Organizations may say they prioritize secure development without backing it with secure practices or a security culture. A global survey of coders about their workplaces found gaps in practice, culture, and training.
The paper proposes PATD, a scheme for privacy-preserving auditing and transparent deduplication of UAV data stored in the cloud. It targets integrity checks that do not reveal file contents to the auditor and that confirm deduplication is performed correctly. The source text is an introduction and related-work discussion and does not state the scheme's results.
OpenAI released Model Spec Evals, a suite that measures how well models follow the OpenAI Model Spec, along with 596 evaluation prompts and open-source evaluation code. Compliance rates were 72% for GPT-4o, 80% for OpenAI o3, 82% for GPT-5 Instant, 89% for GPT-5 Thinking, 84% for GPT-5.3 Instant, and 87% for GPT-5.4 Thinking. The evaluations cover only text-only interactions, and the prompt collection is small relative to the Spec's scope.
Researchers propose MPOF, a privacy-preserving multi-modal object fusion model for connected autonomous vehicles that fuses LiDAR and camera detections using secret sharing. The model detects malicious behavior by a third-party random number generator through sacrificial verification. The authors report that their protocols reduce computational overhead by five orders of magnitude compared with homomorphic-encryption-based random number generation.
This article addresses composite resource allocation problems with general constraints, common in networked systems such as smart grids and multiagent coordination. It proposes two anti-disturbance proximal neural networks: one for structured disturbances, built on the internal model principle, and an observer-based one for unstructured disturbances. Asymptotic convergence of both is established via Lyapunov stability theory, and numerical simulations validate their effectiveness and robustness.
Researchers propose a differentially private accelerated distributed gradient tracking algorithm for distributed aggregative optimization, where each agent's objective depends on an aggregate of all agents' decisions. The method perturbs exchanged information with independent Laplace noise, adds a noise deduction mechanism to keep estimates accurate, and combines dynamic average consensus with heavy-ball momentum. Under strong convexity and Lipschitz-continuous gradients, the authors prove linear convergence in mean-square error, derive explicit suboptimality bounds, and establish that the algorithm satisfies epsilon-DP; numerical simulations are provided.
This article presents a systematic review and meta-analysis of research on GAN-based adversarial attacks against intrusion detection systems (IDSs). The authors empirically test binary and multiclass IDS models under GAN-based attacks, examining the influence of feature selection. They assess IDS performance and sample realism using Fréchet inception distance and Hellinger distance, and identify the most severe GAN architecture and the one that generates the most realistic adversarial samples.
Korbak and colleagues test whether alignment midtraining, which trains a model on fictional documents depicting aligned AI behavior, generalizes to frontier-style models. They replicate the pipeline of an o4-mini-sized model and compare it against misalignment midtraining from Tice et al. The authors report negative early results: the alignment effect fades after reasoning posttraining and does not carry over to more realistic chat and agentic evaluations.
This paper proposes a clean-label backdoor attack on Graph Neural Networks that injects a single trigger node attached to a target node. The attack misclassifies the target and its one-hop neighbors as the target class without changing their features or labels. Across multiple real-world benchmarks and mainstream GNN architectures, the authors report over 95% attack success rate on target nodes and their neighbors in most settings, including under state-of-the-art defenses.
Researchers propose a deepfake face detection method that recasts detection as a Visual Question Answering task for a Visual Language Model. Component-specific prompts direct attention to the eyes, nose and mouth, and a Q-Former module adjusts visual feature focus according to those prompts. The authors report that the method outperforms existing techniques in detection accuracy and robustness, though the source text gives no numerical results.
GDetox is a defense against backdoored graph encoders in graph self-supervised learning (GSSL), the first of its kind per the source. It purifies the encoder through self-supervised distillation without label information, adding adversarial contrastive learning to strengthen the teacher model. Across four node and four graph classification datasets, GDetox reduced the attack success rate to 4% with encoder performance degrading by no more than 2%.
Fix: GDetox: a backdoor defense that purifies a backdoored graph encoder via self-supervised distillation, using adversarial contrastive learning to strengthen the teacher model.
IEEE Xplore (Security & AI Journals)Researchers propose a redundant fully connected layer that can be plugged into existing image classification model backbones to improve adversarial robustness. The layer is trained with a loss function that uses cosine similarity to maximize the difference and diversity among multiple fully connected parts. Tested against 12 representative white-box and black-box attacks on two popular datasets, the approach reportedly gives significant robustness with negligible extra training overhead and hardly degrades clean-sample accuracy.
Fix: The proposed mitigation is the redundant fully connected layer, combined with existing model backbones in a pluggable manner, trained with the cosine-similarity-based loss function described in the source.
IEEE Xplore (Security & AI Journals)This research paper analyzes how backdoor attacks work against semi-supervised learning (SSL), finding that attackers exploit pseudo-labeling to build stronger trigger-target correlations and that triggers can appear in low, medium, or high frequency bands. The authors propose Backdoor Invalidator (BI), a defense combining complementary learning, trigger mix-up, and dual domain filtering. Evaluation against state-of-the-art attacks shows BI significantly reduces average attack success rate while keeping comparable clean-data accuracy.
Fix: BI, a defense framework integrating complementary learning, trigger mix-up, and dual domain filtering, which the authors describe as a plug-in component. Code is available at https://github.com/wxr99/Backdoor_Invalidator4SSL
IEEE Xplore (Security & AI Journals)This paper proposes a Propose-Rectify framework for image manipulation localization that pairs a forensic-adapted LLaVA model with a Forensics Rectification Module and an Enhanced Segmentation Module that adds forensic cues to SAM's image embeddings. The authors report state-of-the-art performance across diverse datasets, with strong robustness and generalization, though the source text gives no numeric results.
Researchers propose MOSA-S2, a multiobjective simulated annealing method that generates text rubbish examples by substituting input words with meaningless stopwords. The method uses importance-based composite perturbation and a grammatically constrained variant, and it outperforms prior approaches across six text datasets and seven neural models. The authors report that models often keep the same prediction, sometimes with higher confidence, on nonsensical text.
FedDOT is a federated learning framework that defends against targeted attacks, including attacks by adversaries who make up more than half of the participants. It combines two detection algorithms, MST-AD (maximum spanning tree-based attacker detection) and density-AD (densest graph-based attacker detection), which use correlation between weight updates and graph theory. Experiments on image classification datasets in non-IID settings report an attack success rate below 10% for single-label flipping, multilabel flipping and backdoor attacks, with an accuracy drop under 2%.