Research
Academic papers, new techniques, benchmarks, and theoretical findings in AI/LLM security.
Academic papers, new techniques, benchmarks, and theoretical findings in AI/LLM security.
11 items
This paper proposes a new way to protect substitution-permutation network (SPN) symmetric encryption, such as AES, in the white-box attack context, where adversaries can see the implementation and control the execution platform. The approach adds secret components to lookup tables to expand and substantially change the encryption's internal states, while the ciphertext stays essentially unchanged. The authors report resistance to existing and unknown attacks in security analysis and good performance across platforms in experiments.
Boaz Barak, Gabriel Wu, Jeremy Chen and Manas Joglekar publish a follow-up to their confessions paper, giving deeper analysis of how training affects confessions and preliminary comparisons to chain-of-thought monitoring. The approach trains a second model output, a confession, rewarded solely for honesty about misbehavior in the main task. The authors hypothesize that honest confessions are the path of least resistance because confessing is easier than sustaining an elaborate lie.
Researchers collaborating with a media house tested whether multimodal AI systems (text and image) could be misused to create and promote misleading news. They generated deceptive news images with five Stable Diffusion variants, including Stable Diffusion 3 and a DPO-tuned model, across the Snopes, Pheme and PolitiFact datasets. They then measured how well the images aligned with news text and whether GPT-4 could classify and rank them compared with human judgment.
Researchers address continual forgetting, the problem of repeatedly removing selected information from a pre-trained vision model while preserving the rest. They propose GS-LoRA, which trains separate LoRA modules on the Feed-Forward Network layers of Transformer blocks for each forgetting task and applies group sparse regularization to select and zero out LoRA groups, and an extension, GS-LoRA++, that adds prototype-based supervision. Experiments on face recognition, object detection and image classification show that specific classes are forgotten with minimal impact on other classes.
BlindU is a machine unlearning method for federated learning that removes a user's data contribution without uploading the erasing data to the server. The user locally produces compressed, privacy-preserving representations through an information bottleneck encoder, and the server unlearns using only those representations and their labels. The authors add a noise-free differential privacy masking step before compression and report better privacy protection and unlearning effectiveness than the best existing privacy-preserving benchmarks.
OpenAI researchers released CoVal, an experimental dataset of crowd-written, prompt-specific rubrics that show why people prefer one model response over another, in addition to which response they chose. CoVal-full keeps the raw, sometimes conflicting criteria, while CoVal-core keeps 4 highly rated, mutually compatible criteria per prompt. The authors state that the rubrics reflect surveyed participants' views, not OpenAI's, and do not represent what all people want from AI.
The paper proposes RODIO, a method that makes deep-learning MRI reconstruction models more robust by using pretrained diffusion models as purifiers. Unlike adversarial training, it avoids a minimax optimization problem and needs only efficient fine-tuning on purified examples. The authors report that it outperforms standalone diffusion-based reconstructors and robustification methods including adversarial training and randomized smoothing across several models, samplers and unseen lesions.
Researchers show that a server adversary in Split Learning (SL) can steal a victim client's functionality, training data and labels. The attack, named SLeak, builds a substitute client that mimics the target's feature extraction using the smashed data and server model. It needs only partial same-domain auxiliary public data and outperforms the prior state-of-the-art method across multiple metrics.
GHAttack is a generative adversarial attack method against heterogeneous graph neural networks (HGNNs). A trained perturbation generator produces a perturbation for each target node in a single forward pass, modifying edges across heterogeneous relations to degrade predictions on target nodes. The authors report high efficiency and effectiveness across ten HGNNs and six datasets.
Researchers show that data augmentation, a standard pre-processing step, can undo the protection of unlearnable examples, raising the accuracy of a model trained on protected data from 21.3% to 66.1%. They propose Armor, a defense framework that uses a non-local module-assisted surrogate model, a surrogate augmentation selection strategy and dynamic step size adjustment to preserve unlearnability under augmentation. Across 4 datasets and 5 augmentation methods, Armor reduces test accuracy on augmented protected samples by as much as 60% more than six baseline defenses.
Fix: Armor is the proposed defense: it adds defensive noise to samples, using a surrogate model and augmentation selection to keep protected data unlearnable under data augmentation. The authors state they will open-source their code upon publication.
IEEE Xplore (Security & AI Journals)This paper reformulates model lineage determination as a question of whether two models' parameters lie within the same local optimum of the loss landscape. The authors propose a method for lineage determination and a task-agnostic, modification-type-agnostic way to measure lineage closeness, based on the mean adversarial distance from data points to decision boundaries and the matching rate of predictions. Experiments report 100% accuracy in lineage determination across a wide range of modification scenarios.