Research
Academic papers, new techniques, benchmarks, and theoretical findings in AI/LLM security.
Academic papers, new techniques, benchmarks, and theoretical findings in AI/LLM security.
19 items
Researchers propose a layered needs-affordances-features framework for designing fair AI hiring systems, with a discovery stage that uses theory to find discrimination mechanisms in digital trace data and a utilization stage that turns them into design interventions. Applied to a dataset of 2,506 applicants' interview responses and historical hiring decisions, they identify language-based discrimination and interview question structure as key mechanisms. Modifying linguistic features and increasing interview structure reduced subgroup disparities, with the strongest gains when both were applied jointly.
Ppin is an anomaly-based intrusion detection system that operates on provenance graphs, which model relationships between system entities. It uses a memory-augmented neural network to model processes against learnable latent categories, then abstracts rare entities into behavior patterns and correlates them by causal dependency to produce alerts that reconstruct an attack narrative. The authors report that Ppin outperforms state-of-the-art systems in node-level detection and produces contextual alerts.
A meta-analysis of 183 independent studies compares informal controls (clan and self-controls) with formal controls in promoting employee information security policy (ISP) compliance. The authors find informal controls are effective and generally more effective than formal controls, with their influence stable across contextual and methodological factors. Formal controls indirectly enhance ISP compliance by shaping perceptions of informal controls through socialization and internalization.
This study examines trust concerns in emerging global hydrogen supply chains linking sources in the Global South and Australia with customers in the Northern Hemisphere. Based on 19 semi-structured interviews with supply chain experts from Germany and Australia, analyzed with the Gioia methodology, the authors identify five affordances for trust building: transparency and traceability, economic risk coordination, scenario planning, knowledge sharing and learning, and communication.
This paper studies security in clustered distributed storage systems (CDSSs), a class of heterogeneous storage systems whose nodes are organized into multiple clusters. It analyzes resiliency against two active adversaries, an omniscient adversary and a limited-knowledge adversary, who can send corrupted data during file reconstruction and repair. The authors derive tight upper bounds on these capacities, achievable through combinatorial coding constructions in repair-by-transfer coding, and these bounds generalize existing results for traditional distributed storage systems.
This research presents a Cross-Layer LSS Vulnerability Characterization Framework for LiDAR subsystems in autonomous vehicles. It applies LiDAR Perception-Pipeline Mechanism Characterization (PMC) to Apollo V8.0's LiDAR perception pipeline, identifying Frame Partition, Object Erasure, and Object Mark as security-critical mechanisms. Validation on a controlled testbed and a real-AV platform shows failures depend on timing, frame-position alignment, frame assembly, and object-lifecycle behavior rather than generic sensor-data corruption alone.
HardFlow is a framework for hard-constrained sampling in flow-matching generative models, which reformulates constraint enforcement as a trajectory optimization problem solved with numerical optimal control and model predictive control techniques. The method steers sampling so constraints hold precisely at the terminal time, rather than projecting the entire path onto the constraint manifold as prior approaches do. Experiments in robotics planning, PDE boundary control and text-guided image editing report that HardFlow substantially outperforms existing methods on constraint satisfaction and sample quality.
The v2026.06 release of the AI Sec Watch content adds the techniques Steal Web Session Cookie, Use Alternate Authentication Material: Web Session Cookie, and AI Service Web Interface, and updates the LLM Jailbreak technique. It also updates the mitigations Generative AI Guardrails, Generative AI Guidelines, and AI Telemetry Logging, and adds six case studies, including Storm-2139 Azure OpenAI Guardrail Bypass and EchoLeak zero-click prompt injection against M365 Copilot.
A systematic literature review analyzed 50 peer-reviewed literature reviews from 2019 to 2024 on digital marketing, branding, and the metaverse, drawn from Scopus and AISeL. Most contributions were theoretical (68%), and the review identified four core ethical dimensions: privacy, transparency, authenticity, and fairness. It argues these challenges require interdisciplinary responses from marketers, platform providers, policymakers, and consumers.
Researchers propose BUUAS, a Belief-updating and User-aware Security framework for scheduling prompt injection detection in LLM service deployments across edge-cloud networks. The framework combines a Bayesian-inspired belief update method with a belief-weighted contextual multi-armed bandit mechanism to prioritize high-risk requests. Experiments report that it outperforms state-of-the-art methods in detection accuracy, network throughput and resource efficiency under fluctuating user loads and varying malicious user ratios.
This research introduces LanEvil++, the first benchmark for testing lane perception robustness in autonomous driving under environmental illusions such as shadows and tire marks. Built with the CARLA simulator, it covers 14 illusion types across 94 3D scenes, and evaluations show lane detection models lose 5.27% accuracy and 10.49% F1-score, with shadows the most disruptive. The authors also propose the Multimodal Illusion Defense Approach (MIDA), which uses hard examples to improve robustness by 4.23% on lane detection models and 3.82% on ADVLMs.
Fix: The authors propose the Multimodal Illusion Defense Approach (MIDA), which uses hard examples to improve illusion resistance, boosting robustness by 4.23% on LD models and 3.82% on ADVLMs.
IEEE Xplore (Security & AI Journals)Fix: By matching trust-building techniques as functional enablers, we provide specific advice on how these concerns can be mitigated.
AIS eLibrary (Journal of AIS, CAIS, etc.)In early 2024, Google's Gemini image generation model produced racially inaccurate and historically inconsistent images, such as non-White figures in Nazi-era contexts. The source attributes the failure to a "diversity injection" mechanism meant to counter data bias, which lacked contextual safeguards. The backlash reportedly cost Google $90 billion in market value.
This paper introduces algorithmic fragility, the persistent instability of AI systems in organizational settings, framed as a structural condition rather than a temporary defect. It proposes stabilization work, the distributed labor organizations use to absorb and manage algorithmic breakdowns, and identifies three micro-level practices: buffering, reframing, and patching.
Researchers show that poisoning a retrieval-augmented generation (RAG) system can amplify bias in an LLM's outputs, even for gender-neutral queries. Their Bias Retrieval and Reward Attack (BRRA) framework generates adversarial documents using multi-objective reward functions, manipulates retrieval with subspace projection, and uses a cyclic feedback mechanism, with experiments on several mainstream models showing significant bias increases. The paper also explores a dual-stage defense mechanism to mitigate the attack.
Fix: The source mentions a dual-stage defense mechanism that it says can effectively mitigate the impacts of the attack, but it does not describe its specifics, configuration or implementation.
IEEE Xplore (Security & AI Journals)TraceNeRF is a method for tracing the training datasets of Neural Radiance Fields (NeRF) models, not the models themselves. It embeds owner-specific binary messages into the datasets through a hybrid frequency-spatial framework that combines a learnable spatially aware mask with discrete cosine transform, plus a density-aware module and a trainable band selector. The authors report that it outperforms baseline methods in availability, effectiveness, and robustness.
Researchers propose PromptFishing, an active method for identifying LLM-driven social accounts by embedding machine-readable prompts in ordinary topics that induce a target hallucinated response. The method uses a gradient-guided two-stage beam search to generate adversarial suffixes, first pushing responses off topic and then aligning them with the target. On Twitter data, it reports a true positive rate above 0.85 with a false positive rate below 0.01.
This paper proposes Self-Evolving Neuro-Symbolic Autonomy (SENSA), a neuro-symbolic approach for Cyber-Physical Systems that combines neural-symbolic bidirectional self-alignment, causal-driven self-evolution and hierarchical resilience planning. The authors claim the method adds causal reasoning and robustness to adversarial noise in open-world settings. They report theoretical proofs and empirical results on eight benchmark datasets, where it outperforms state-of-the-art baselines.
The paper SemAder: Evading LLM-Based Binary Code Analysis via Structure-Semantics Joint Induction was published in ACM Transactions on Privacy and Security, Volume 29, Issue 3, pages 1-34, in August 2026. The source text consists only of this publication metadata and does not describe the method or findings.
Researchers present EthicAR, a hierarchical Safe Reinforcement Learning framework for autonomous driving that adds ethics-aware cost signals combining collision probability and harm severity. The method adds risk-sensitive prioritized experience replay and Temporal Cost Aggregation, which aligns learning with tail-risk measures such as Conditional Value-at-Risk. In closed-loop simulations on the Waymo Open Dataset across 75 scenarios and five random seeds, it reduced collision rates by 20 to 45% versus baselines while keeping task success and comfort within 5 to 10% of baselines.