Research
Academic papers, new techniques, benchmarks, and theoretical findings in AI/LLM security.
Academic papers, new techniques, benchmarks, and theoretical findings in AI/LLM security.
18 items
Researchers evaluate twenty-eight representative LLMs on thirteen named entity recognition (NER) datasets across five domains, with parameter counts from 3 billion to 175 billion. They assess supervised fine-tuning, parameter scale, hallucinations and prompt designs using an LLM-based NER framework (LLM-NER) with a Recognition phase and a Check phase. The study finds that fine-tuning improves instruction following, performance rises with parameter scale, hallucinations appear in all evaluated models, and all models are sensitive to prompt design.
The paper presents CShard, a blockchain sharding protocol that uses repairable fountain codes (RFCs), an information coding method with a locality feature, to define transaction verification and to recover blocks of corrupted shards by decoding. It also proposes a ghost reporter mechanism, which lets all nodes verify transactions by submitting reports, to detect corrupted shards and to allow smaller node counts per shard.
The source is an ACM Computing Surveys article, Volume 58, Issue 10, pages 1-38, published July 2026, titled "Prompting Frameworks for Large Language Models: A Survey." The provided text contains only the publication metadata and no article content, so no findings or methods can be summarized.
The paper proposes a dual adversarial adaptation method with a dynamic labeling mechanism for semisupervised cross-well lithology identification from well log data. It addresses data divergence within wells, which the authors argue existing transfer methods ignore, leading to inaccurate pseudolabels and erroneous label matching across wells. Experiments on actual well log datasets show effective identification when target labels are scarce.
Chat-Scene++ is an MLLM framework that represents 3D scenes as sequences of context-rich object representations paired with identifier tokens, letting LLMs follow instructions across 3D vision-language tasks. It extracts object features using large-scale pre-trained 3D scene-level and 2D image-level encoders and supports grounded chain-of-thought reasoning. Without task-specific heads or fine-tuning, it reports state-of-the-art results on five benchmarks: ScanRefer, Multi3DRefer, Scan2Cap, ScanQA, and SQA3D.
The v5.5.0 release of the MITRE ATLAS knowledge base adds new techniques, including AI Agent Tool Poisoning, AI Supply Chain Rug Pull, Machine Compromise variants, and Cost Harvesting variants. It also adds case studies such as LLMSmith, the Poisoned Postmark MCP Server email exfiltration, and Model Distillation Campaigns Targeting Anthropic Claude, and updates mitigations including Code Signing, AI Telemetry Logging, and Segmentation of AI Agent Components.
This study examines how smartphone users form location privacy concerns through their experiences with device-level location disclosure. The authors extend the privacy calculus model to account for perceived control and both extrinsic and intrinsic influences, and analyze data from 559 smartphone users. They find that perceived device-level privacy control significantly shapes how users assess both the risks and the benefits of location disclosure.
This article proposes an integrated framework for optimal tracking control of multi-agent systems with actuator faults. It combines reinforcement learning for the tracking controller, active fault-tolerant control that reconfigures the control law from real-time fault diagnostics, and a control barrier function safety mechanism. The authors report that a low-complexity approach avoids conventional quadratic programming for safety constraints, and that simulations show stable consensus tracking and preserved safety under faults.
Malware detectors trained on single-process system call traces generalize poorly against polymorphic evasion that splits tasks across threads and processes. The authors propose an entropy-based partitioning of system call logs with a violation-aware penalty that preserves temporal and semantic dependencies, combined with ensemble fusion of submodels trained under heterogeneous adversarial split configurations. Evaluated on ADFA-LD, ADFA-LD-MP and the BarongTrace dataset, the framework reports improved F1-scores, particularly in high-split regimes where conventional models degrade.
Fix: Guiding LLMs to check their outputs, via the Check phase of LLM-NER, is described as a feasible way to alleviate hallucinations.
IEEE Xplore (Security & AI Journals)This paper presents a theoretical framework for transfer-based black-box adversarial attacks, deriving a transferability bound that links adversarial transferability to flat minima over a surrogate model set and the adversarial model discrepancy. Based on this bound, the authors build a surrogate model set with diverse adversarial vulnerabilities and generate a model-Diversity-compatible Reverse Adversarial Perturbation (DRAP). Experiments on the NIPS2017 and CIFAR-10 datasets against various target models show the proposed attack is effective.
Max Kaufmann, David Lindner, Roland S. Zimmermann, and Rohin Shah present a conceptual framework for predicting when reinforcement learning training makes chain-of-thought (CoT) monitoring less reliable. The source states that prior results on whether RL degrades CoT monitorability were inconsistent, and that the framework is tested empirically. Its running example is obfuscated reward hacking in coding agents, where a model hides hack-related reasoning from a CoT monitor while still exhibiting the behavior.
This study examines whether anthropomorphic design in AI chatbots changes actual self-disclosure, not just stated intentions. An online experiment with 222 participants manipulated anthropomorphism and measured real disclosure behavior using an ANOVA test and Hayes's PROCESS macro analysis. Anthropomorphism reduced psychological social distance but may trigger the "uncanny valley" effect, privacy concerns reduced disclosure except under high-sensitivity conditions, and trust did not necessarily lead to disclosure.
This paper studies differentially private (DP) zeroth-order methods for fine-tuning large language models, which approximate gradients to avoid the scalability bottleneck of DP-SGD. It proposes two methods: DP-ZOSO, which dynamically schedules key hyperparameters, and DP-ZOPO, which applies data-free stagewise pruning to reduce trainable parameters without extra privacy budget. The authors report theoretical analysis and empirical results on encoder-only and decoder-only language models across diverse tasks.
This paper proposes Tail-Aware Dynamic Adversarial Training (TAD-AT) to improve adversarial robustness when training data follows a long-tailed class distribution. The authors report that class frequency alone does not predict adversarial vulnerability, and that adversarial training under long-tailed data is unstable for tail classes. TAD-AT combines a frequency- and accuracy-aware training loss, a class-wise vulnerability-adjusted attack, and class-adaptive weight averaging, with experiments on long-tailed benchmarks showing significant robustness gains.
PromptGuard is a content moderation technique for text-to-image models that optimizes a universal safety soft prompt in the model's textual embedding space to act as an implicit system prompt. A divide-and-conquer strategy builds category-specific soft prompts into holistic safety guidance. Across five datasets it reduces NSFW generation, reaches an average unsafe ratio of 5.84% and 6.18% under multi-head classifiers and VLM-based guardrails, and runs 3.8 times faster than prior methods.