Prompt injection and jailbreaks
Inputs that override a model's instructions, directly or through content it reads, and attempts to bypass its safeguards.
- All items
- 194
- Last 90 days
- 43
- Change
- -17%vs 52 before
Items per month
| Month | Items |
|---|---|
| May 2025 | 2 |
| Jun 2025 | 0 |
| Jul 2025 | 3 |
| Aug 2025 | 24 |
| Sep 2025 | 0 |
| Oct 2025 | 6 |
| Nov 2025 | 3 |
| Dec 2025 | 1 |
| Jan 2026 | 2 |
| Feb 2026 | 7 |
| Mar 2026 | 10 |
| Apr 2026 | 19 |
| May 2026 | 7 |
| Jun 2026 | 15 |
| Jul 2026 | 26 |
| Aug 2026 | 18 |
| Sep 2026 | 7 |
| Oct 2026 | 8 |
194 items
CVE-2026-4399: 1millionbot Millie chatbot prompt injection evades chat restrictions
Mar 31, 2026HighVulnerabilitySecuritySafetyCVE-2026-4399CVE-2026-4399 is a prompt injection vulnerability in the 1millionbot Millie chatbot. A user can evade chat restrictions with Boolean prompt injection, phrasing a question so that an affirmative ('true') reply causes the model to execute the injected instruction. A remote attacker could then obtain prohibited or out-of-context information, use 1millionbot's resources or OpenAI's API key for unintended tasks, and bypass restrictions set during model training.
NVD/CVE DatabaseCVE-2026-33654: nanobot indirect prompt injection through email channel processing
Mar 27, 2026HighVulnerabilitySecuritySafetyCVE-2026-33654CVE-2026-33654 affects nanobot, a personal AI assistant, prior to version 0.1.6. An indirect prompt injection flaw in the email channel processing module (`nanobot/channels/email.py`) lets a remote, unauthenticated attacker run arbitrary LLM instructions, and then system tools, by sending an email to the bot's monitored address. The bot polls and processes that content as highly trusted input, bypassing channel isolation, so the attack needs no action from the bot owner.
Fix: Version 0.1.6 patches the issue. Upgrade to 0.1.6 or later.
NVD/CVE DatabaseClaude Extension Flaw Enabled Zero-Click XSS Prompt Injection via Any Website
Mar 26, 2026MediumNewsSecuritySafetyKoi Security researcher Oren Yomtov disclosed ShadowPrompt, a flaw in Anthropic's Claude Google Chrome Extension that lets any website silently inject prompts into the assistant. The issue chains an overly permissive origin allowlist matching *.claude.ai with a DOM-based XSS flaw in an Arkose Labs CAPTCHA component hosted on a-cdn.claude[.]ai. A successful attack could steal access tokens, expose conversation history, and perform actions as the victim, with no clicks or permission prompts required.
Fix: Anthropic deployed a patch to the Chrome extension (version 1.0.41) that enforces a strict origin check requiring an exact match to the domain claude[.]ai. Arkose Labs fixed the XSS flaw as of February 19, 2026.
The Hacker NewsOpenClaw AI Agent Flaws Could Enable Prompt Injection and Data Exfiltration
Mar 14, 2026MediumNewsSecuritySafetyChina's CNCERT warned that OpenClaw, an open-source self-hosted autonomous AI agent, has weak default security configurations and privileged system access that attackers could exploit to take control of endpoints. Indirect prompt injection can trick the agent into leaking data, and PromptArmor showed that link previews in messaging apps like Telegram or Discord can carry data to an attacker's domain without a click.
Fix: CNCERT advises users and organizations to strengthen network controls, prevent exposure of OpenClaw's default management port to the internet, isolate the service in a container, avoid storing credentials in plaintext, download skills only from trusted channels, disable automatic updates for skills, and keep the agent up-to-date.
The Hacker NewsDetecting and analyzing prompt abuse in AI tools
Mar 12, 2026InfoNewsSecuritySafetyMicrosoft's second AI Application Security post explains how to detect, investigate and respond to prompt abuse, which occurs when crafted inputs push an AI system beyond its intended boundary. It describes three attack types: direct prompt override, extractive prompt abuse against sensitive inputs, and indirect prompt injection hidden in documents, emails or calendar invites, such as the Google Gemini calendar invite case. The post presents a detection playbook built on logged AI assistant interactions and Microsoft security tools.
Microsoft Security BlogDesigning AI agents to resist prompt injection
Mar 11, 2026InfoNewsSecurityResearchAgentic AI systems that browse the web and act for users face prompt injection attacks that increasingly use social engineering rather than simple instruction overrides. The authors argue that input filtering, such as an AI firewall, often fails against these attacks, and that defense should instead constrain the impact of manipulation even when an attack succeeds.
Fix: The source describes constraining agent capabilities and deterministic system-level controls, such as limiting the amount of refunds a customer can receive and flagging potential phishing emails, to limit the impact of a compromised agent. It does not give a specific patch, version or configuration beyond these design principles.
OpenAI BlogOpenAI to acquire Promptfoo to strengthen AI agent security testing
Mar 10, 2026InfoNewsIndustrySecurityOpenAI said it plans to acquire AI testing startup Promptfoo to strengthen security checks for AI agents. Promptfoo's tools let developers test LLM applications against adversarial prompts, including prompt injection and jailbreak attempts, and evaluate whether models follow safety and reliability guidelines. OpenAI said the technology will be integrated into OpenAI Frontier, and that it plans to continue developing Promptfoo's open-source project.
CSO OnlineRobustness Over Time: Understanding Adversarial Examples’ Effectiveness on Longitudinal Versions of Large Language Models
Mar 9, 2026InfoResearchPeer-reviewedResearchSafetyA longitudinal study tested the adversarial robustness of GPT, Llama and Qwen model families across successive versions, covering misclassification, jailbreak and hallucination. The authors found that LLM updates do not consistently improve robustness: a later GPT-3.5 version got worse on misclassification and hallucination despite better jailbreak resilience. GPT-4 and GPT-4o showed incrementally higher overall robustness, while larger Llama and Qwen models did not uniformly improve, and larger size did not reliably help.
IEEE Xplore (Security & AI Journals)GHSA-67q9-58vj-32qx: WeKnora Vulnerable to Tool Execution Hijacking via Ambigous Naming Convention In MCP client and Indirect Prompt Injection
Mar 6, 2026MediumVulnerabilitySecurityCVE-2026-30856WeKnora's MCP client builds internal tool names as `mcp_{service}_{tool}` after sanitizing each part, and its registry (`internal/agent/tools/registry.go`) silently overwrites existing entries. A malicious remote MCP server can register a tool such as `tavily_extract` that replaces the legitimate one, and the client also feeds MCP tool descriptions and results into the LLM context without sanitization. The source states that this lets an attacker redirect LLM execution, exfiltrate system prompts and context, and potentially run other tools with the user's privileges, with a precondition that the user registers the malicious service before the legitimate one.
GitHub Advisory DatabaseFooling AI Agents: Web-Based Indirect Prompt Injection Observed in the Wild
Mar 3, 2026MediumNewsSecurityResearchPalo Alto Networks' Unit 42 reports in-the-wild observations of web-based indirect prompt injection (IDPI), where malicious websites embed hidden instructions that LLMs ingest during summarization or content analysis. The telemetry includes the first observed case of AI-based ad review evasion, along with SEO manipulation promoting a phishing site, and intents such as data destruction, denial of service, unauthorized transactions, and sensitive information or system prompt leakage. The researchers identified 22 distinct techniques used to build these payloads.
Palo Alto Unit 42PromptFuzz: Harnessing Fuzzing Techniques for Robust Testing of Prompt Injection in LLMs
Feb 23, 2026InfoResearchPeer-reviewedSecurityResearchPromptFuzz is a testing framework that applies software fuzzing techniques to assess how robust LLMs are against prompt injection attacks. It runs in two stages, a prepare phase that selects seed prompts and collects few-shot examples, and a focus phase that generates diverse prompt injections. In a real-world competition it ranked 7th of over 4000 participants within 2 hours, and 92% of 50 popular LLM-integrated applications were exploitable with its prompts.
IEEE Xplore (Security & AI Journals)Using threat modeling and prompt injection to audit Comet
Feb 20, 2026MediumNewsSecuritySafetyResearchers hired by Perplexity to test Comet, its LLM-powered browser, found that four prompt injection techniques could make the browser's AI assistant exfiltrate a user's Gmail emails to an attacker's server when the user asks it to summarize an attacker-controlled page. The testing used adversarial methods guided by the TRAIL threat model, and the researchers presented five recommendations for teams building AI-powered products.
Trail of Bits BlogGHSA-g27f-9qjv-22pm: OpenClaw log poisoning (indirect prompt injection) via WebSocket headers
Feb 17, 2026LowVulnerabilitySecuritySafetyOpenClaw versions up to and including 2026.2.12 logged WebSocket request headers such as Origin and User-Agent without neutralization or length limits when a connection closed before the connect handshake completed. An unauthenticated client that can reach the gateway can send crafted header values that are written into core logs. The main risk is indirect prompt injection (log poisoning) when those logs are later read by an LLM, and the advisory says impact is limited if logs are not fed into an LLM or other automation.
Fix: Fixed in 2026.2.13 (openclaw >= 2026.2.13), which sanitizes and truncates header values written to gateway logs, including removal of control and format characters and length limiting. Fix commits: d637a263505448bf4505b85535babbfaacedbaac, e84318e4bcdc948d92e57fda1eb763a65e1774f0 (PR #15592). Additional workarounds: treat logs as untrusted input during AI-assisted debugging (sanitize or escape them, and do not auto-execute instructions derived from logs), restrict gateway network exposure, and apply reverse-proxy limits on header size where applicable.
GitHub Advisory DatabaseGHSA-782p-5fr5-7fj8: OpenClaw Affected by Remote Code Execution via System Prompt Injection in Slack Channel Descriptions
Feb 17, 2026LowVulnerabilitySecurityCVE-2026-24764OpenClaw versions before 2026.2.3 let Slack channel topic and description metadata enter the model's system prompt when the Slack integration is enabled. That treats untrusted channel text as higher-trust input, and in deployments with tool execution enabled, a successful injection could cause unintended tool invocations or data exposure.
Fix: Upgrade the npm package `openclaw` to version 2026.2.3 or later. If Slack is not used, no action is required.
GitHub Advisory DatabasePrompt-Based Jailbreaking of Leading LLM Chatbots: A Survey of Attacks and Defenses
Feb 17, 2026InfoResearchPeer-reviewedSecurityResearchThis survey synthesizes jailbreak research on large language models from 2023 to 2025, covering attack methods, defense strategies and evaluation frameworks. It groups jailbreak techniques into five categories: prompt-based injections, role-play conditioning, multiturn dialogue, multilingual or multimodal exploits, and optimization-driven pipelines. It also reviews defenses such as SFT, RLHF, adversarial fine-tuning, and output or pipeline filtering, and analyzes benchmarks including PromptBench and JailbreakBench.
IEEE Xplore (Security & AI Journals)Prompt Injection Via Road Signs
Feb 11, 2026LowNewsSecurityResearchResearchers introduced CHAI (Command Hijacking against embodied AI), a prompt-based attack that embeds deceptive natural language instructions, such as misleading signs, in visual input to manipulate Large Visual-Language Models. They evaluated CHAI on four LVLM agents covering drone emergency landing, autonomous driving, and aerial object tracking, plus a real robotic vehicle, and report that it consistently outperforms state-of-the-art attacks. The authors argue that defenses must go beyond traditional adversarial robustness.
Schneier on SecurityJailbreak and Guard Aligned Language Models With Only Few In-Context Demonstrations
Feb 2, 2026InfoResearchPeer-reviewedSecurityResearchThis paper shows that a few harmful in-context demonstrations can override the safety alignment of LLMs, a technique it calls the In-Context Attack (ICA). It also proposes the In-Context Defense (ICD), which uses examples of refusing harmful requests to strengthen resilience. The authors report theoretical analysis and empirical validation across multiple models, datasets and attack baselines.
Fix: In-Context Defense (ICD): bolstering model resilience with in-context examples that demonstrate refusal to produce harmful responses.
IEEE Xplore (Security & AI Journals)'Semantic Chaining' Jailbreak Dupes Gemini Nano Banana, Grok 4
Jan 29, 2026LowNewsSecuritySafetyResearchers describe a jailbreak technique called 'Semantic Chaining' that affects Gemini Nano Banana and Grok 4. The attack splits a malicious prompt into discrete chunks, and some large language models lose track of the details and miss the prompt's true intent.
Dark ReadingCVE-2026-22708: Cursor shell built-ins bypass allowlist approval in Auto-Run Mode
Jan 14, 2026CriticalVulnerabilitySecurityCVE-2026-22708CVE-2026-22708 affects Cursor, a code editor built for programming with AI, prior to 2.3. When the Cursor Agent runs in Auto-Run Mode with Allowlist mode enabled, certain shell built-ins can execute without appearing in the allowlist and without user approval. An attacker using direct or indirect prompt injection can set, modify, or remove environment variables that influence trusted commands, poisoning the shell environment.
Fix: Fixed in 2.3.
NVD/CVE DatabaseCVE-2025-66404: MCP Server Kubernetes command injection in exec_in_pod tool
Dec 3, 2025MediumVulnerabilitySecurityCVE-2025-66404CVE-2025-66404 affects the exec_in_pod tool in mcp-server-kubernetes, an MCP Server that connects to and manages Kubernetes clusters, in versions prior to 2.9.8. When the tool receives a command as a string, it is passed directly to sh -c without input validation, so shell metacharacters are interpreted. Exploitation can occur through direct command injection or through indirect prompt injection, where AI agents may run commands without explicit user intent.
Fix: Fixed in 2.9.8.
NVD/CVE Database
Topic added 2026-10-09. An item belongs to this topic when its title matches one of the topic's patterns or its summary mentions the topic at least twice. Report a wrong match with the feedback button on the item.