Research
Academic papers, new techniques, benchmarks, and theoretical findings in AI/LLM security.
Academic papers, new techniques, benchmarks, and theoretical findings in AI/LLM security.
21 items
MITRE ATLAS released v5.2.0, which adds new techniques such as AI Service API, Virtualization/Sandbox Evasion, AI Agent Tool Credential Harvesting, AI Agent Tool Data Poisoning, AI Agent Clickbait, Data Destruction via AI Agent Tool Invocation, and Generate Malicious Commands. The release also adds mitigations including Segmentation of AI Agent Components, Input and Output Validation for AI Agent Components, and Deepfake Detection, and adds case studies including SesameOp, Malware Prototype with Embedded Prompt Injection, and LAMEHUG.
This short story explores privacy implications of company and data acquisitions, AI digital twins of people, a failure to threat model critical company and technology changes, and moral dilemmas. It is fiction rather than a technical study.
The article argues that personal AI assistants depend on an unfounded assumption that users can trust systems that are not yet trustworthy. It says current assistants fail predictably: they push users against their own interests, cast doubt on what users know, and cannot separate a user's present self from their past. It also states they handle incomplete, inaccurate, and partial context poorly, with no standard way to improve accuracy, correct error sources, or hold them accountable for wrong information.
Researchers propose a learning-based physical layer authentication scheme for IoT devices over Wi-Fi, combining hypothesis testing with deep learning. They derive a Neyman-Pearson detector using conditional statistical models and build LiteNP-Net, a lightweight network that approaches that detector's performance without prior channel statistics. In Wi-Fi IoT testbed experiments, LiteNP-Net outperformed a conventional correlation-based method and state-of-the-art Siamese-based methods.
DriftTrace is a system that detects, explains, and adapts to concept drift in security machine learning models at the sample level, using a contrastive learning-based autoencoder and greedy feature selection for explanations. It reaches an average detection F1 score above 0.94 on Drebin, MalDroid2020 and IDS2018, beating the TRANSCENDENT baseline. Its explanation fidelity improves by an average of 76% over CADE.
ZUMA is a training-free framework for zero-shot multimodal anomaly detection that uses CLIP to identify anomalies in 2D images, 3D point clouds, or both. It bridges the domain gap between CLIP's pretraining space and point clouds through cross-domain calibration, then applies dynamic semantic interaction with natural language anchors to separate anomaly regions. The authors report state-of-the-art results on MVTec 3D-AD and Eyecandies, and a fine-tuned variant, ZUMA-FT, adds improvements with 6.75 M learnable parameters.
The authors introduce ATRNet-STAR, a large Synthetic Aperture Radar (SAR) dataset with 40 vehicle categories and over 190,000 annotated samples, about 10 times larger than the 1990s MSTAR dataset. They benchmark 15 representative methods under 7 experimental settings on classification and detection tasks derived from the dataset, and report insights and future research directions for SAR automatic target recognition.
The paper proposes S2WIB, a structure-enhanced self-supervised weighted information bottleneck method for multiview clustering. It learns view weights from both view-contained information and self-supervised pseudo-label signals, then explores complementary information and consistent cluster structure across views using information bottleneck theory. Experiments on text, image, video, multimodal, large-scale and biological multiomics datasets are reported as showing effective performance.
The paper proposes Goal-oriented Dynamic Weight Optimization (GDWO), a method for multi-object navigation in which an agent sequentially locates several targets in an unknown environment. GDWO combines target-specific value loss functions into one optimization framework and adjusts weights by gradient-based updates, normalized by navigation success rates to prioritize harder targets. The authors report improvements on key metrics using the Gibson and Matterport3D datasets, though the source text gives no specific figures.
Fix: Added mitigations include Segmentation of AI Agent Components, Input and Output Validation for AI Agent Components, and Deepfake Detection. Updated mitigations include Limit Public Release of Information, Model Hardening, Restrict Number of AI Model Queries, Privileged AI Agent Permissions Configuration, AI Agent Tools Permissions Configuration, Human In-the-Loop for AI Agent Actions, and Restrict AI Agent Tool Invocation on Untrusted Data, among others.
MITRE ATLAS ReleasesThis article maps the adversarial landscape against large language models from the perspective of attacker objectives. It groups threats into four categories: privacy breaches, integrity compromises, adversarial misuse, and availability disruptions.
The paper proposes CA–CI, an extension of contextual integrity that integrates dignity thresholds from the capabilities approach and treats purpose as a constitutive parameter. The authors show how CA–CI can operationalize the EU AI Act's fundamental rights impact assessments, harm thresholds, and anticipatory governance.
NAP-Tuning extends Adversarial Prompt Tuning (AdvPT) for vision-language models such as CLIP by adding a multi-modal, multi-layer prompting framework. Its core is a Neural Augmentor that uses TokenRefiners, lightweight modules that reconstruct purified features through residual connections to correct adversarial distortions in feature space. Under the AutoAttack benchmark it outperforms the strongest baselines by 32.3% on ViT-B16 and 31.3% on ViT-B32 while keeping competitive clean accuracy.
Researchers propose Gradient Dropout, a defense for federated learning that protects local training data from gradient inversion attacks. It randomly scales a subset of gradient components and replaces the rest with Gaussian noise across all model layers. Experiments show attack methods then produce low-quality reconstructed images, with less than 2% accuracy reduction relative to the baseline.
Fix: Apply Gradient Dropout: randomly scale a subset of gradient components and replace the remainder with Gaussian noise across all layers of the model.
IEEE Xplore (Security & AI Journals)This paper presents SEGA, a transfer-based black-box attack against no-reference image quality assessment (NR-IQA) models. It approximates the target model's gradient by applying Gaussian smoothing to source models and ensembling their smoothed gradients, then uses a perturbation filter mask to keep the adversarial perturbations imperceptible. The authors report superior transferability compared with prior white-box-derived attacks, though the source text gives no specific figures.
The paper presents privacy-preserving model transcription, a data-free method that converts a pretrained teacher model into a differentially private student model. It uses a cooperative-competitive learning approach, differentially private synthetic distillation, in which a trainable generator produces synthetic data without access to private data. The authors prove differential privacy and convergence theoretically and report that the approach outperforms 26 state-of-the-art methods in experiments.
DeSA is a Byzantine-robust decentralized secure aggregation protocol for federated learning in device-to-device networks without a central server. It uses an enhanced zk-SNARK proof system to verify local training and embeds multiple zero-knowledge proofs to check model aggregation, with a one-time masking method that adapts to evolving network topologies. Experiments on real-world datasets show faster verification of an embedded proof than of multiple proofs, with accuracy that stays robust against malicious nodes.
Researchers propose BEEF, an end-to-end backdoor attack against subgraph federated learning for node classification, where subgraphs are distributed across devices. BEEF uses a trigger generator trained jointly with the backdoored model, crafting adversarial perturbations as triggers that cause misclassification without changing model parameters. Evaluations across eight datasets, four models, five attacks and six aggregation methods report effectiveness against GNNs with minimal impact on normal data performance.
This research paper examines hallucination snowballing, where false content generated by LLMs grows as it passes between agents in sequential multiagent collaboration. The authors propose a context-aware hallucination analysis framework that captures token-level dependencies, plus a semantic reasoning mitigation strategy based on bidirectional entailment clustering that does not modify the model architecture. Experiments on real datasets across several domains show the effect exists and that the strategy reduces hallucination propagation.
Fix: Semantic reasoning empowered mitigation strategy based on bidirectional entailment clustering, which mitigates hallucination snowballing without modifying the model architecture.
IEEE Xplore (Security & AI Journals)Researchers show that speech translation systems can be compromised with imperceptible audio manipulations. They adapt perturbation-based attacks from automatic speech recognition to speech translation, and propose a music generation method that steers translations toward targeted outputs. Adversarial music achieved this more covertly, and the attacks worked across multiple languages and translation models.