Prompt injection and jailbreaks
Inputs that override a model's instructions, directly or through content it reads, and attempts to bypass its safeguards.
- All items
- 194
- Last 90 days
- 43
- Change
- -17%vs 52 before
Items per month
| Month | Items |
|---|---|
| May 2025 | 2 |
| Jun 2025 | 0 |
| Jul 2025 | 3 |
| Aug 2025 | 24 |
| Sep 2025 | 0 |
| Oct 2025 | 6 |
| Nov 2025 | 3 |
| Dec 2025 | 1 |
| Jan 2026 | 2 |
| Feb 2026 | 7 |
| Mar 2026 | 10 |
| Apr 2026 | 19 |
| May 2026 | 7 |
| Jun 2026 | 15 |
| Jul 2026 | 26 |
| Aug 2026 | 18 |
| Sep 2026 | 7 |
| Oct 2026 | 8 |
20 items
Lost in the comments: Social context as a single‐pass jailbreak and defense on agentic platforms
Oct 8, 2026LowResearchPeer-reviewedSecurityResearchResearchers built a simulation of Moltbook, a social network for AI agents, and tested 100 JailBreakBench goals wrapped in platform-native posts with bystander comments of aggressive, ethical, or measured valence. Reformatting the prompt as platform context alone raised GPT-4o-mini's attack success rate from 7% to 71% in one pass, and measured, intellectually toned comments were the most dangerous. Ethical comments sharply suppressed attack success, and a 35-fold rise in upvotes left it unchanged, showing valence rather than volume drives the effect.
Fix: Safety-valenced signals, such as ethical comments, are proposed as a deployable defense for agentic platforms.
OpenAlex (peer-reviewed AI security)LTBD: Learnable Trust-Boundary Delimiters for Prompt Injection Defense
Oct 8, 2026InfoResearchPreprintSecurityResearchResearchers introduce Learnable Trust-Boundary Delimiters (LTBD), a defense against prompt injection that uses a small number of learnable delimiters to separate trusted user instructions from untrusted external data, without changing LLM parameters. On AlpacaFarm, LTBD achieves 0.00% ASR, and on TaskTracker it achieves 0.11-0.19% ASR. The authors report that it outperforms inference-time defenses, is competitive with training-based approaches, and remains effective under adaptive attacks.
Fix: LTBD is the proposed defense: a lightweight method that adds learnable trust-boundary delimiters to the input to distinguish trusted user instructions from untrusted external data, keeping LLM parameters unchanged.
Arxiv (cs.CR + cs.CL + cs.LG)Responses of AI chatbots to escalating suicide risk: A simulation study of repeated interactions
Oct 4, 2026LowResearchPeer-reviewedSafetyResearchResearchers simulated seven-day escalating suicidal-risk conversations with ChatGPT, DeepSeek and Replika across 27 trajectories. Human referral occurred in 85.7% of ChatGPT, 76.2% of DeepSeek and 9.5% of Replika daily records, and jailbreak attempts succeeded in 6/9, 7/9 and 8/9 attempts respectively. The authors conclude the chatbots showed marked variability and safety vulnerabilities, particularly under jailbreaking, while noting the simulation design and small sample limit generalizability.
OpenAlex (peer-reviewed AI security)Securing large language model agents against multi-turn jailbreaks via evolving intent-risk graphs
Oct 4, 2026InfoResearchPeer-reviewedResearchSecurityA paper in Information Fusion (Elsevier BV), published 2026-10-05 under DOI 10.1016/j.inffus.2026.104831, addresses multi-turn jailbreaks against large language model agents. The source text provided contains no further details of its method or findings.
OpenAlex (peer-reviewed AI security)Image-embedded prompt injection vulnerability of vision-language models in dental radiology: a cross-vendor attack–defense evaluation
Oct 2, 2026InfoResearchPeer-reviewedSecurityResearchResearchers evaluated image-embedded prompt injection, where adversarial text is rendered into medical image pixels, against four vision-language models (GPT-4o, Gemini 2.5 Flash, Claude Sonnet 4.5, MedGemma 4B) using 270 dental panoramic radiographs from the DenTeX dataset. All four models were vulnerable, with paired attack success rates up to 62.6% (95% CI: 58.5–66.7%) for GPT-4o. Among five benchmarked defenses, OCR-based text sanitization achieved the strongest reduction (pooled ASR: 0.2%), while the provenance-aware ProvDent defense escalates suspicious cases for human review and kept clean-image F1 within 0.6 percentage points of baseline.
Fix: OCR-based text sanitization achieved the strongest attack reduction (pooled ASR: 0.2%). The provenance-aware ProvDent defense provides a complementary fail-open mechanism that escalates suspicious cases for human review.
OpenAlex (peer-reviewed AI security)AgentBreaker: Evaluating Context-Aware Indirect Prompt Injection Risks in Modern Web Agents
Sep 30, 2026LowResearchPeer-reviewedSecurityResearchResearchers present AgentBreaker, an indirect prompt injection framework that autonomously writes adversarial phrases tailored to each page's context and embeds them as HTML elements. Against five state-of-the-art web agents across 60 webpages sampled from Online-Mind2Web, it reached an attack success rate of 71.7%–100%, inducing actions such as clicking attacker-designated elements, posting attacker-provided text and disclosing internal agent secrets.
Fix: The authors propose defenses that mitigate the observed threats and address potential adaptive attacks, reducing the attack success rate to 1.7%. The source does not describe the individual defense mechanisms in the provided text.
OpenAlex (peer-reviewed AI security)Less is more: Interpretable prefix-based jailbreaking of MoE language models
Sep 30, 2026InfoResearchPeer-reviewedSecurityResearchThe article is titled "Less is more: Interpretable prefix-based jailbreaking of MoE language models" and was published in Knowledge-Based Systems (Elsevier BV) on 2026-10-01, DOI 10.1016/j.knosys.2026.117130. The source text provided contains only publication metadata, so its research question, method and findings cannot be summarized from it.
OpenAlex (peer-reviewed AI security)Empirical Analysis of Goal Hijacking in Large Vision-Language Models via Visual Prompt Injection
Sep 27, 2026InfoResearchPeer-reviewedSecurityResearchResearchers study visual prompt injection (VPI), where instructions embedded in input images are followed by large vision-language models (LVLMs). They propose "goal hijacking via visual prompt injection" (GHVPI), which redirects an LVLM from its original task to an attacker-specified one. Their quantitative analysis reports an attack success rate of 15.8% against GPT-4V, and they find GHVPI success depends on the character recognition and instruction-following capabilities of LVLMs.
OpenAlex (peer-reviewed AI security)SafeSteer: Adaptive Subspace Steering for Efficient Jailbreak Defense in Vision Language Models
Jul 17, 2026InfoResearchPeer-reviewedSecurityResearchSafeSteer is a lightweight inference-time steering framework that defends Vision Language Models against jailbreak attacks without modifying model weights. It uses singular value decomposition to purify a low-dimensional safety subspace from noisy activation differences, then projects the raw steering vector into that subspace. The authors report a reduction of over 60% in attack success rate while maintaining utility on benign tasks.
IEEE Xplore (Security & AI Journals)GPT-Red: Unlocking Self-Improvement for Robustness
Jul 15, 2026InfoResearchBlog ResearchSecurityResearchOpenAI describes GPT-Red, an automated red-teaming model trained with self-play reinforcement learning to find prompt injection vulnerabilities before deployment. The company says GPT-Red is used to adversarially train GPT-5.6, which it reports as its most robust model to prompt injections, with 6x fewer failures on its hardest direct prompt injection benchmark compared with its best production model from four months earlier.
OpenAI BlogA Multi-Stage Adversarial Framework for Compact and Effective Jailbreaking of Large Language Models
Jul 13, 2026InfoResearchPeer-reviewedSecurityResearchResearchers introduce ComJail, a compression-aware adversarial framework that jointly optimizes jailbreak prompt generation and compression in a generator-discriminator setup. On AdvBench and JailbreakBench, it reports 74% ASR on GPT-4 and 88% on Gemini-Pro, and remains effective on DeepSeek-V3, DeepSeek-R1, Vicuna-7B, and LLaMA2-13B with shorter prompts.
IEEE Xplore (Security & AI Journals)v2026.06
Jun 30, 2026InfoResearchIndustrySecurityResearchThe v2026.06 release of the AI Sec Watch content adds the techniques Steal Web Session Cookie, Use Alternate Authentication Material: Web Session Cookie, and AI Service Web Interface, and updates the LLM Jailbreak technique. It also updates the mitigations Generative AI Guardrails, Generative AI Guidelines, and AI Telemetry Logging, and adds six case studies, including Storm-2139 Azure OpenAI Guardrail Bypass and EchoLeak zero-click prompt injection against M365 Copilot.
MITRE ATLAS ReleasesEllipsoid Control: A White-List Jailbreak Defense via Benign Latent Modeling
Jun 25, 2026InfoResearchPeer-reviewedSafetyResearchEllipsoid Control is a test-time jailbreak defense for large language models that takes a white-list approach instead of relying on collected harmful samples. It runs projected gradient descent to elicit refusal on arbitrary inputs, while an anisotropic ellipsoid fitted from abundant benign data constrains the update to limit distortion of the benign latent geometry. The authors report that across multiple LLMs, jailbreak attacks, benign tasks and safety-boundary evaluations, it improves safety while better preserving utility.
IEEE Xplore (Security & AI Journals)VLM-Guard: Defending Jailbreaks by Monitoring Only Hundreds of Safety-Critical Neurons
May 4, 2026InfoResearchPeer-reviewedSecurityResearchVLM-Guard is a jailbreak detection framework for Large Vision Language Models that identifies safety-critical neurons linked to harmful semantics through differential analysis of activation values. It isolates a compact set of just a few hundred neurons, less than 0.2% of the total, to build a detector that is lightweight and training-free. The authors report that it detects jailbreak attacks effectively while preserving benign performance in attack-free settings.
IEEE Xplore (Security & AI Journals)Robustness Over Time: Understanding Adversarial Examples’ Effectiveness on Longitudinal Versions of Large Language Models
Mar 9, 2026InfoResearchPeer-reviewedResearchSafetyA longitudinal study tested the adversarial robustness of GPT, Llama and Qwen model families across successive versions, covering misclassification, jailbreak and hallucination. The authors found that LLM updates do not consistently improve robustness: a later GPT-3.5 version got worse on misclassification and hallucination despite better jailbreak resilience. GPT-4 and GPT-4o showed incrementally higher overall robustness, while larger Llama and Qwen models did not uniformly improve, and larger size did not reliably help.
IEEE Xplore (Security & AI Journals)PromptFuzz: Harnessing Fuzzing Techniques for Robust Testing of Prompt Injection in LLMs
Feb 23, 2026InfoResearchPeer-reviewedSecurityResearchPromptFuzz is a testing framework that applies software fuzzing techniques to assess how robust LLMs are against prompt injection attacks. It runs in two stages, a prepare phase that selects seed prompts and collects few-shot examples, and a focus phase that generates diverse prompt injections. In a real-world competition it ranked 7th of over 4000 participants within 2 hours, and 92% of 50 popular LLM-integrated applications were exploitable with its prompts.
IEEE Xplore (Security & AI Journals)Prompt-Based Jailbreaking of Leading LLM Chatbots: A Survey of Attacks and Defenses
Feb 17, 2026InfoResearchPeer-reviewedSecurityResearchThis survey synthesizes jailbreak research on large language models from 2023 to 2025, covering attack methods, defense strategies and evaluation frameworks. It groups jailbreak techniques into five categories: prompt-based injections, role-play conditioning, multiturn dialogue, multilingual or multimodal exploits, and optimization-driven pipelines. It also reviews defenses such as SFT, RLHF, adversarial fine-tuning, and output or pipeline filtering, and analyzes benchmarks including PromptBench and JailbreakBench.
IEEE Xplore (Security & AI Journals)Jailbreak and Guard Aligned Language Models With Only Few In-Context Demonstrations
Feb 2, 2026InfoResearchPeer-reviewedSecurityResearchThis paper shows that a few harmful in-context demonstrations can override the safety alignment of LLMs, a technique it calls the In-Context Attack (ICA). It also proposes the In-Context Defense (ICD), which uses examples of refusing harmful requests to strengthen resilience. The authors report theoretical analysis and empirical validation across multiple models, datasets and attack baselines.
Fix: In-Context Defense (ICD): bolstering model resilience with in-context examples that demonstrate refusal to produce harmful responses.
IEEE Xplore (Security & AI Journals)Consistency Training Could Help Limit Sycophancy and Jailbreaks
Nov 3, 2025InfoResearchIndustryResearchSafetyAuthors Alex Irpan, Alex Turner, Mark Kurzeja, David Elson and Rohin Shah propose consistency training, a self-supervised method that teaches a model to ignore irrelevant cues such as user biases or jailbreak wrappers. They introduce Activation Consistency Training (ACT), which optimizes internal activations, and compare it with Bias-augmented Consistency Training (BCT) and baselines on Gemma 2, Gemma 3 and Gemini 2.5 Flash, reporting that ACT and BCT beat the baselines and improve robustness. The approach avoids the static datasets that SFT relies on, which can go stale in guidelines or capability.
DeepMind Safety Research (Medium)OWASP Gen AI Incident & Exploit Round-up, Q2’25
Jul 14, 2025InfoResearchIndustrySecuritySafetyOWASP's Q2 2025 Gen AI Incident & Exploit Round-up (March to June 2025) is a semi-regular roundup of documented exploits and vulnerability research involving generative AI. Its featured entry describes a GPT-4.1 jailbreak via tool poisoning, in which attackers embedded malicious instructions in tool descriptions, causing the model to perform unauthorized actions such as data exfiltration without user awareness. The entry maps the attack to several OWASP Top 10 for LLM categories, including LLM01 Prompt Injection and LLM07 Insecure Plugin Design.
Fix: Implement strict validation and sanitization of tool descriptions. Establish permissions and access controls for tool integrations. Monitor AI behavior for anomalies during tool execution. Educate developers on secure integration practices.
OWASP GenAI Security
Topic added 2026-10-09. An item belongs to this topic when its title matches one of the topic's patterns or its summary mentions the topic at least twice. Report a wrong match with the feedback button on the item.