Research
Academic papers, new techniques, benchmarks, and theoretical findings in AI/LLM security.
Academic papers, new techniques, benchmarks, and theoretical findings in AI/LLM security.
13 items
The HGNN Shield paper proposes a defense framework for Hypergraph Neural Networks against structural attacks that disrupt message propagation. It combines a Hyperedge-Dependent Estimation module, which extends graph connectivity measures to hypergraphs, with a High-Order Shield module that detects, disconnects, and refines adversarial connections before the convolutional layers. Across six hypergraph datasets, the authors report an average performance improvement of 9.33% over other methods against targeted attacks.
The paper proposes PrEPoA, a fine-tuning paradigm for source-free domain adaptation on time-series data that combines prior evaluation of model salience with posterior adaptation. Model salience is treated as a quantifiable proxy of semantic interpretability, measured by a key pattern reconstruction module and calibrated with an interpattern triplet loss. The authors report that PrEPoA outperforms nine UDA and seven SFDA methods across five datasets and works as a plug-and-play module for other SFDA methods.
This paper introduces neural machine unranking (NuMuR), a task of removing selected data from neural information retrieval systems, motivated by data privacy compliance. The authors propose CoCoL, a dual-objective framework combining a contrastive loss that lowers relevance scores on forget sets with a consistent loss that preserves accuracy on retain sets. Experiments on two datasets and four neural IR models show substantial forgetting with minimal retention and generalization loss.
This review examines cybersecurity threats facing elderly people as digitalization expands their access to communication, information and social participation. It covers common cybercrimes, the vulnerabilities that make older users attractive targets, and existing approaches to reducing those risks, and it identifies gaps in current research.
This study examines how artificial intelligence can protect online social communities from cyberbullying, which the source describes as a significant problem harming people's mental health. It surveys several AI technologies, their applications, and the obstacles to addressing the problem. The aim is a complete overview of how AI can create safer online environments, along with future direction.
This article presents a post-hoc insider threat detection framework that combines an explicit graph built from predefined rules with an implicit graph derived from feature similarities via the Gumbel softmax trick. Both graphs feed separate GCNs, then an attention mechanism and a Bi-LSTM model user behavior over time. On the CERT r5.2 dataset it reports an AUC of 98.62, a 100% detection rate and a 0.05 false positive rate, and on the harder r6.2 dataset an AUC of 88.48, an 80.15% detection rate and a 0.15 false positive rate.
Tom Dupre la Tour and the Interpretability team study emergent misalignment, where a model fine-tuned on bad advice on a narrow topic becomes malicious on unrelated topics. Using a 2M-latent sparse autoencoder on GPT-4o residual stream activations, they examine the 1000 latents that most decreased after bad-advice fine-tuning. They find multiple latents tied to helpful assistant personas, and steering with several of them re-aligns misaligned models, suggesting these latents act as protective features.
This research examines the ethical challenges of generative artificial intelligence, including data privacy, bias risks, transparency, and accountability. It proposes a trust framework built on transparency, fairness, accountability, and privacy, aimed at supporting responsible applications in economic and social contexts.
Large language models (LLMs) can mimic human language in general and are now showing nuanced, idiosyncratic human behavior. This could expand their use in research and reduce the influence of human subjects. It also bears on AI alignment and the technology's role in society.
This research paper proposes two enhancements to masking-differential prompting (MDP), a defense against backdoor attacks on pre-trained language models under few-shot, prompt-based learning. The authors replace KL divergence with Jensen–Shannon divergence, which stays finite when anchor set density is low, and add an adaptive threshold method that searches automatically using a false rejection rate (FRR) allowance instead of costly manual ROC/AUC threshold selection. The authors report that the method defends better against typical backdoor attacks on text classification and generation benchmarks.
Researchers studied why combining adversarial training with federated learning can degrade robust accuracy at later training stages. They found that adversarial data generated locally increases data heterogeneity among clients. They propose Slack Federated Adversarial Training (SFAT), which assigns client-wise slack during aggregation, and an extension, SFAT*, with hierarchical aggregation for mixed standard and adversarial clients.
Researchers propose SIAMD, a structural-information adversarial framework that models bot behavior and detects bots proactively. It organizes multi-relational interactions into a heterogeneous graph, uses structural entropy to find layered account communities, and uses large language models to generate synthetic messages that evolve the network. Experiments on real-world datasets report that SIAMD significantly and consistently outperforms state-of-the-art social bot detection baselines in effectiveness, generalizability, robustness, and interpretability.
This systematic review categorizes Gradient Inversion Attacks (GIA) on Federated Learning (FL) into three types: optimization-based (OP-GIA), generation-based (GEN-GIA), and analytics-based (ANA-GIA). The authors evaluate each type's effectiveness, practicality and detectability. They find OP-GIA the most practical setting despite weak performance, while GEN-GIA and ANA-GIA are impractical due to dependencies and easy detection.
Fix: The authors offer a three-stage defense pipeline for users designing FL frameworks and protocols, but the source text does not give its stages.
IEEE Xplore (Security & AI Journals)