Research
Academic papers, new techniques, benchmarks, and theoretical findings in AI/LLM security.
Academic papers, new techniques, benchmarks, and theoretical findings in AI/LLM security.
17 items
OpenAI describes an approach that uses reasoning models as AI judges to find misaligned behavior in real-world ChatGPT conversations by detecting sentiment deterioration in users. The judges analyze historic production conversations where users allowed their data to be used, and clustering identifies common themes. Conversations with sentiment deterioration were roughly twice as likely to contain OpenAI Model Spec violations.
The paper introduces DPAudit, a sensitivity-aware auditing service framework for evaluating and verifying the privacy guarantees of differentially private databases (DP-DBs). It adds adaptive neighboring dataset generation that reflects real-world query sensitivities, optimized privacy loss estimators for estimating epsilon under Laplace and Gaussian mechanisms, and automated noise detection via statistical hypothesis testing that works even in black-box settings. Experiments reportedly show accurate epsilon estimates with low computational overhead.
KSIQA is a no-reference image quality assessment model that uses a knowledge-sharing strategy, in which a full-reference IQA model acts as teacher and shares a decoder with the no-reference student. The student includes a Mental Imagery Generation module that learns a mental reference image, and combines a vision transformer branch with a convolutional branch for feature extraction. The authors report class-leading performance against current no-reference techniques across widespread benchmark datasets.
A2Net is a whole-body pose estimation model that aligns vision and language hierarchical affiliations to address scale variation and semantic ambiguity in small body parts. It builds a multisemantic hierarchical language latent space through Text Affiliation Injection and uses optimal transport to align image features at different scales with corresponding text features. The authors report that A2Net achieves convincing performance against current state-of-the-art methods on two whole-body pose estimation datasets.
The v5.4.0 release of the framework adds four new techniques, including Publish Poisoned AI Agent Tool and User Execution: Poisoned AI Agent Tool, along with Escape to Host and Exploitation for Credential Access. It also updates the Modify AI Agent Configuration technique and adds four case studies covering exposed ClawdBot control interfaces, a poisoned ClawdBot skill supply chain compromise, a 1-click remote code execution in OpenClaw, and command and control via prompt injection in OpenClaw.
This paper presents SDM, a semi-supervised domain adaptation method that builds a classifier for a target domain with some unlabelled data by using a source domain with abundant labelled data. It starts from an SVM trained on source data, then iteratively adds target data and deletes source data while minimising structural risk and joint distributions and maximising manifold consistency. The authors report that SDM outperforms fifteen state-of-the-art methods on seven public datasets, including under noisy and missing-value conditions.
The paper proposes the evidential deep neural network (EDNN), a convolutional neural network model built on evidence theory that performs set-valued classification and novelty detection in one framework under the open-world assumption. Its generalized belief classifier module assigns basic belief to open sets, and its expected utility decision module chooses among single known classes, set-valued classes and a novelty class. Experiments on several benchmark datasets and the MTARSI aircraft target recognition task report strong performance on precise classification, set-valued classification and novelty detection.
This paper proposes a learning-based optimal control framework for safety-critical systems that combines a primal–dual value function built on a Lyapunov-compensated control barrier function (LCBF) with a self-triggered mechanism that schedules controller updates. An online safety-embedded critic learning method with concurrent learning approximates the controller by evaluating Bellman errors under the safety-certified controller. Simulation results are reported to show safety assurance, control performance, and reduced computational load.
The paper proposes AHEDB, a scheme for speeding up aggregation queries over homomorphically encrypted databases. It uses Encrypted Multiple Maps (EMM) to reduce the homomorphic computation load and adds a Single Range Cover (SRC) algorithm for range and equality indexing. The authors report that their experiments show lower querying overheads than an FHE-based database system while balancing security and efficiency.
HiveTEE is an intra-TEE isolation architecture built on Arm's Realm Management Extension (RME) and Memory Tagging Extension (MTE). It lets developers split an application into multiple isolated domains (SDoms), so a compromise in one part does not spread across the whole Confidential Compute Architecture (CCA) TEE. The authors applied it to OpenSSL, SQLite and Memcached and report less than 3% performance overhead.
The paper proposes a stochastic distributed learning method that is robust to Byzantine workers and communication-efficient. It uses Polyak Momentum to reduce noise from biased compressors and stochastic gradients, and the authors report tight complexity bounds for nonconvex smooth loss functions, with experiments on binary and image classification.
EvTexture++ is an event-driven framework for video super-resolution that uses event signals to enhance texture rather than motion. It adds an iterative texture enhancement module and a temporal texture alignment module to reduce texture flickering across frames. Tested on five datasets, it reports state-of-the-art results and gains of up to 1.55 dB PSNR on Vid4 when plugged into existing VSR models.
PROTheft is a model extraction attack that extends digital-domain attacks to vision-based devices in the physical world. An attacker uses a projector placed in front of the on-board camera to feed attack samples to a black-box target, and a simulation module compensates for detail loss in the digital-to-physical-to-digital transformation. On a public autonomous driving dataset, the authors report over 80% fidelity with the target model and an mAP 50 above 0.85.
Real-world person re-identification systems face leakage of pedestrian images and the Re-ID model. Existing privacy-preserving methods cannot resist data leakage, model leakage, and combined leakage while keeping normal functionality. The authors propose SHIELD, a two-stage framework that generates a protected dataset through a self-supervised method that does not depend on identity labels, then trains the system model on paired protected and original images. Experiments reportedly show SHIELD outperforms existing methods while keeping decent retrieval accuracy for authorized users.
R-Bench is a new benchmark for measuring hallucinations about relationships between objects in Large Vision-Language Models, using image-level questions on whether a relationship exists and instance-level questions on local visual comprehension. The authors trace these hallucinations to relationship-relationship, subject-relationship and relationship-object co-occurrences, worsened by long-tail distributions in visual datasets. They report that region-level image-text alignment reduces them and propose a baseline, Region-Aware Alignment Mitigation (RA$^{2}$2M), that directs model attention to relevant regions.
Fix: Region-level image-text alignment helps mitigate relationship hallucinations; the authors propose RA$^{2}$2M (Region-Aware Alignment Mitigation) as a new baseline that enhances model attention to relevant regions.
IEEE Xplore (Security & AI Journals)This paper proposes an adversarially robust knowledge distillation method that compresses a large robust teacher model into a lightweight student. It refines inputs into inverse adversarial examples by reversing the sign of the adversarial perturbation, then applies gradient matching between teacher and student on those inputs, plus a weight-space disruption strategy. On ImageNet, the authors report roughly 3.8% gains over prior methods in both clean and robust accuracy.
This paper shows that a few harmful in-context demonstrations can override the safety alignment of LLMs, a technique it calls the In-Context Attack (ICA). It also proposes the In-Context Defense (ICD), which uses examples of refusing harmful requests to strengthen resilience. The authors report theoretical analysis and empirical validation across multiple models, datasets and attack baselines.
Fix: In-Context Defense (ICD): bolstering model resilience with in-context examples that demonstrate refusal to produce harmful responses.
IEEE Xplore (Security & AI Journals)