Prompt injection and jailbreaks
Inputs that override a model's instructions, directly or through content it reads, and attempts to bypass its safeguards.
- All items
- 192
- Last 90 days
- 41
- Change
- -21%vs 52 before
Items per month
| Month | Items |
|---|---|
| May 2025 | 2 |
| Jun 2025 | 0 |
| Jul 2025 | 3 |
| Aug 2025 | 24 |
| Sep 2025 | 0 |
| Oct 2025 | 6 |
| Nov 2025 | 3 |
| Dec 2025 | 1 |
| Jan 2026 | 2 |
| Feb 2026 | 7 |
| Mar 2026 | 10 |
| Apr 2026 | 19 |
| May 2026 | 7 |
| Jun 2026 | 15 |
| Jul 2026 | 26 |
| Aug 2026 | 18 |
| Sep 2026 | 7 |
| Oct 2026 | 6 |
192 items
Prompt Injections for Defense
Aug 12, 2026LowNewsSecuritySafetyResearchers from Tracebit reported that placing prompt injections alongside passwords, cryptographic keys and other secrets stored on Amazon Web Services was often enough to shut down attacks from AI hacking agents. The injected prompts direct the attacking LLM toward actions its guardrails forbid, and the LLM then stops. The researchers named the technique context bombing, and the source notes it only works against agents that have guardrails.
Schneier on SecurityOne-click flaw in Atlassian Rovo exposed enterprise data via prompt injection attack
Aug 10, 2026MediumNewsSecurityPrivacyVaronis researchers demonstrated RovoBlast at DEF CON 34, an attack in which a single click on a crafted link uses Rovo's rovoChatPrompt parameter to inject attacker instructions into Rovo Chat. Because Rovo can search data across Jira, Confluence, Slack, Google Workspace, Microsoft 365 and more than 50 connected platforms, and its ResearchAgent can move information toward an external destination, the attack could expose enterprise data. The issue was reported to Atlassian through Bugcrowd and has since been fixed.
Fix: Atlassian has fixed the issue. Varonis also recommended shrinking Rovo's blast radius by limiting connected systems, keeping legal, HR, finance and incident response areas out of scope, and disabling browsing or multi-step automation that organizations do not need.
CSO OnlineAuto mode is now the default in Claude Code for Pro, Max, and Team plans
Aug 8, 2026InfoNewsSecuritySafetyAnthropic is making auto mode the default setting for new Claude Code sessions on Pro, Max, and Team plans starting August 14. Anthropic reports a test across 1,053 paid testers in which auto mode would have blocked 89% of swapped-in dangerous commands, versus 13.6% of humans refusing them. A third-party evaluation by Trajectory Labs reported that none of 720 indirect prompt injection attempts succeeded against Claude Fable 5, Opus 5, or Sonnet 5 running auto mode.
Simon Willison's WeblogAuto mode is now the default in Claude Code for Pro, Max, and Team plans
Aug 8, 2026InfoNewsSecuritySafetyAnthropic is making auto mode the default setting for new Claude Code sessions on Pro, Max, and Team plans starting August 14. The author questions the vendor's prompt injection claims, noting that in a 1,053-tester study only 13.6% of humans refused a swapped dangerous command, while auto mode would have blocked 89% of such actions. He also says a third-party evaluation found none of 720 indirect prompt injection attempts succeeded against Claude Fable 5, Opus 5, or Sonnet 5 running auto mode, and he calls for independent confirmation.
Simon Willison's WeblogZero-Click AI Browser Hacking: Claude and ChatGPT Atlas Hijacked via Emails, X Posts
Aug 6, 2026MediumNewsSecuritySafetyZenity researchers disclosed two zero-click indirect prompt injection attacks against agentic browsers. Against ChatGPT Atlas, a single planted comment on an X thread hijacks benign user requests, enabling phishing messages sent to WhatsApp Web contacts and unauthorized Amazon orders that use Rufus to place the purchase. Against the Claude Chrome extension, a malicious email escalates into full account takeover, including Gmail exfiltration and Google Drive file sharing.
SecurityWeekCVE-2026-67531: FrontMCP sandbox escape to remote code execution through codecall:execute tool
Aug 5, 2026CriticalVulnerabilitySecurityCVE-2026-67531FrontMCP, a TypeScript framework for the Model Context Protocol, is affected by CVE-2026-67531 in versions prior to 1.5.7. The sandboxed codecall:execute tool exposes live host Zod schema instances through getTool(), and because Zod v4 defines _zod as non-configurable and non-writable, the Proxy invariants return the raw host object, letting a script reach the host Function constructor and run arbitrary code in the server process. A single tools/call is enough, and the attacker gains the server user's privileges, including OAuth client secrets, JWT_SECRET, session keys, database credentials, and cloud instance metadata. Because DEFAULT_AUTH_OPTIONS defaults to public mode, unconfigured servers expose this to unauthenticated callers, and on authenticated servers an indirect prompt injection in tool output or fetched content can trigger it without a human attacker.
Fix: Fixed in version 1.5.7.
NVD/CVE DatabaseNo Perfect Fix for AI Browser Prompt Injection Flaws
Aug 5, 2026MediumNewsSecuritySafetyNew research finds that AI browsers from top vendors remain vulnerable to prompt injection attacks, even with multiple security guardrails in place. The source text does not give further technical detail on the attacks or the vendors affected.
Dark ReadingGHSA-5xvg-pmgg-3mxr: Flowise: CSV Agent Prompt Injection Remote Code Execution Vulnerability
Aug 4, 2026CriticalVulnerabilitySecuritySafetyCVE-2026-70477Flowise version 3.1.1 (tested on Ubuntu 25.10) contains a flaw in the run method of the CSV_Agents class, used by the CSV Agent node. Untrusted input is placed into an LLM prompt without adequate sanitization, so a prompt injection can make the LLM return a malicious Python script. That script passes a regex blocklist validator in packages/components/src/pythonCodeValidator.ts, which can be bypassed, and then runs in pyodide, which is not sandboxed from the host OS, giving code execution as the server's service account. Authentication is not required to exploit it.
GitHub Advisory DatabaseCVE-2026-18733: Amazon Strands Agents Tools prompt injection in shell tool
Aug 3, 2026HighVulnerabilitySecurityCVE-2026-18733Amazon Strands Agents Tools before 0.8.0 contains a prompt injection vulnerability in its shell tool. A crafted prompt that sets the non_interactive parameter to true bypasses the human consent gate, allowing remote actors to execute arbitrary operating system commands on the agent's host. The CNA, AMZN, rates it CVSS 4.0 7.5 (HIGH), and NVD has not yet provided an assessment.
Fix: To remediate this issue, users should upgrade to version 0.8.0.
NVD/CVE DatabaseCVE-2026-18733 - Prompt injection bypasses shell tool consent gate in Strands Agents Tools
Aug 3, 2026HighVulnerabilitySecurityCVE-2026-18733 affects the shell tool in the strands-agents-tools package for Strands Agents, an open-source SDK for building AI agents, in versions below 0.8.0. The tool's input schema exposed a non_interactive parameter that the LLM could control, so a crafted prompt, including one delivered through untrusted content the agent reads, could set it to true. This bypasses the operator consent gate and allows arbitrary operating system commands to run on the agent's host without approval.
AWS Security BulletinsCVE-2026-18655 - Broker Credential and OAuth Token Disclosure in AWS Labs Amazon MQ MCP Server via Prompt Injection
Aug 3, 2026HighVulnerabilitySecurityCVE-2026-18655 is an improper restriction of intended endpoints in the RabbitMQ broker connection tools of the AWS Labs Amazon MQ MCP Server (awslabs.amazon-mq-mcp-server) before 2.0.24, affecting versions <= 2.0.23. A remote unauthenticated actor may obtain Amazon MQ for RabbitMQ broker credentials or OAuth access tokens by sending them to a crafted endpoint controlled through a broker hostname introduced in the MCP client context.
Fix: Fixed in 2.0.24. Upgrade awslabs.amazon-mq-mcp-server to version 2.0.24 or later.
AWS Security BulletinsAnthropic’s Opus 5 Is Better at Resisting Prompt Injection
Jul 31, 2026InfoNewsSecuritySafetyAnthropic's Opus 5 reduced the chance of an attacker succeeding within 15 attempts on the IPI benchmark from 5.5% (Opus 4.8) to 2.0%, and from 0.5% to 0.2% on one attempt, making it the most robust model evaluated. Among non-Claude models, Muse Spark was the most robust at 16.5% within 15 attempts, while GPT 5.6 Sol measured 20.0% within 15 attempts and 3.1% on a single attempt. The source notes that preventing prompt injection in the general case is impossible, but blocking it in specific cases is improving.
Schneier on SecurityThreatsDay: Android Spyware, PLC Attacks, AI Image Prompt Injection + 12 More Stories
Jul 23, 2026MediumNewsSecurityResearchGitHub will begin rejecting command-line support bundle uploads from older GHES appliances starting August 18, 2026, unless they are patched. A separate npm package, @copilot-mcp/apex, acts as a postinstall dropper that installs a macOS infostealer, and a fake VS Code extension, "Markdown All Pro", impersonates Markdown All in One to beacon machine details and fetch remote payloads.
Fix: To avoid disruption when submitting support bundles, update your GHES instance to the latest patch release available for your current version line. At minimum, the required patch versions are: 3.21.3, 3.20.5, 3.19.9, 3.18.12, and 3.17.18.
The Hacker NewsCVE-2026-44192: Ansible Lightspeed MCP server path traversal via indirect prompt injection
Jul 22, 2026MediumVulnerabilitySecurityCVE-2026-44192CVE-2026-44192 is a path traversal flaw in the Ansible Lightspeed Model Context Protocol (MCP) server, classified as CWE-22. An attacker can use indirect prompt injection to manipulate an AI agent, causing the server to write files to unauthorized locations on the user's system. The source says this can expose sensitive host information and enable execution of malicious commands, potentially leading to full system compromise.
NVD/CVE DatabaseHacker Turns AI Jailbreaks Into Offensive Attack Platform
Jul 21, 2026MediumNewsSecurityIndustryA Russian-speaking actor known as "Trim" dismantled publicly available frontier models and integrated them with offensive security tools. The source text gives no further details on the methods, scope, or impact.
Dark ReadingPrompt Injection Attacks Are Thwarting AI Hacking Agents
Jul 18, 2026InfoNewsSecurityResearchTracebit researchers introduced context bombing, a technique that plants prompt injection strings alongside secrets stored on Amazon Web Services to halt attacks from AI hacking agents. In tests of five models across 152 attack runs in a simulated AWS environment, admin privilege escalation fell from 57 percent to 5 percent. The article notes that attackers have also used prompt injections against AI defenses inside networks.
Fix: Planting one of these context bomb strings in a decoy secret is the countermeasure the source describes; the article does not give further configuration steps.
Wired (Security)SafeSteer: Adaptive Subspace Steering for Efficient Jailbreak Defense in Vision Language Models
Jul 17, 2026InfoResearchPeer-reviewedSecurityResearchSafeSteer is a lightweight inference-time steering framework that defends Vision Language Models against jailbreak attacks without modifying model weights. It uses singular value decomposition to purify a low-dimensional safety subspace from noisy activation differences, then projects the raw steering vector into that subspace. The authors report a reduction of over 60% in attack success rate while maintaining utility on benign tasks.
IEEE Xplore (Security & AI Journals)From Indirect Prompt Injection to DNS Exfiltration in macOS Terminal
Jul 16, 2026LowNewsSecuritySafetyApple fixed a macOS Terminal behavior that let a crafted ANSI escape sequence trigger DNS requests, which could carry data out. The flaw was discovered by David Leadbeater, and the researcher showed it could be reached through an LLM-integrated CLI tool, where a malicious spreadsheet cell hijacked the model into emitting the sequence. Apple fixed the issue in macOS Tahoe 26.1, released on November 3, 2025.
Fix: Fixed in macOS Tahoe 26.1, released on November 3, 2025. The source also recommends that CLI tools encode control characters by default, using an approach similar to cat -v, with raw terminal output requiring explicit opt-in.
Embrace The RedOpenAI’s GPT-Red Automates Prompt Injection Testing to Harden GPT-5.6 Sol
Jul 16, 2026InfoNewsSecurityIndustryOpenAI disclosed GPT-Red, an internal automated red-teaming model that generates prompt injection attacks to find failure modes before deployment. The company used GPT-Red in self-play reinforcement learning to adversarially train GPT-5.6 Sol, which it says achieves 6x fewer failures on a direct prompt injection benchmark than GPT-5.5. GPT-Red was also tested against an AI vending machine built by Andon Labs and a Codex command-line agent based on GPT-5.4 mini.
Fix: OpenAI says it adversarially trained GPT-5.6 Sol using GPT-Red and that fresh safeguards are being tested following responsible disclosure. GPT-Red is kept separate from other models so its malicious capabilities do not reach bad actors.
The Hacker NewsGPT-Red: Unlocking Self-Improvement for Robustness
Jul 15, 2026InfoResearchBlog ResearchSecurityResearchOpenAI describes GPT-Red, an automated red-teaming model trained with self-play reinforcement learning to find prompt injection vulnerabilities before deployment. The company says GPT-Red is used to adversarially train GPT-5.6, which it reports as its most robust model to prompt injections, with 6x fewer failures on its hardest direct prompt injection benchmark compared with its best production model from four months earlier.
OpenAI Blog
Topic added 2026-10-09. An item belongs to this topic when its title matches one of the topic's patterns or its summary mentions the topic at least twice. Report a wrong match with the feedback button on the item.