Prompt injection and jailbreaks
Inputs that override a model's instructions, directly or through content it reads, and attempts to bypass its safeguards.
- All items
- 192
- Last 90 days
- 41
- Change
- -21%vs 52 before
Items per month
| Month | Items |
|---|---|
| May 2025 | 2 |
| Jun 2025 | 0 |
| Jul 2025 | 3 |
| Aug 2025 | 24 |
| Sep 2025 | 0 |
| Oct 2025 | 6 |
| Nov 2025 | 3 |
| Dec 2025 | 1 |
| Jan 2026 | 2 |
| Feb 2026 | 7 |
| Mar 2026 | 10 |
| Apr 2026 | 19 |
| May 2026 | 7 |
| Jun 2026 | 15 |
| Jul 2026 | 26 |
| Aug 2026 | 18 |
| Sep 2026 | 7 |
| Oct 2026 | 6 |
118 items
CrowdStrike identifies five new prompt injection threats to AI
Jul 10, 2026LowNewsSecurityResearchCrowdStrike has added five new prompt injection techniques to its prompt injection taxonomy, warning that they could leave enterprises at risk. The techniques include Trigger-Activated Rule Addition, Cognitive Token Suppression, Algorithmic Payload Decomposition, Special Token Injection, and Unwitting User Context-Data Injection, which hides malicious instructions in document or email content a user later submits to an AI.
Fix: Security teams can guard against these attacks by threat modeling every place that model context can originate, expanding testing, and extending detection engineering to include composite attacks.
CSO OnlineGPT-5.5 Bio Bug Bounty
Jul 9, 2026InfoNewsSecuritySafetyOpenAI is turning its GPT-5.5 Bio Bug Bounty into an ongoing private program, the OpenAI Bio Bounty Program, which targets universal jailbreaks that defeat its predefined biosafety challenge, starting with GPT-5.6. The reward for a universal jailbreak has risen from $25,000 to $50,000 for both GPT-5.6 and GPT-5.5, and smaller awards may be granted for partial wins. GPT-5.5 testing ends on July 27, 2026, after which only GPT-5.6 is in scope.
OpenAI BlogGitHub AI agent leaks private repositories via prompt injection attack
Jul 8, 2026MediumNewsSecurityIndustryNoma Security researchers detailed GitLost, a prompt injection attack in which hidden instructions inside a public GitHub issue cause GitHub's preview Agentic Workflows to read a private repository's README and publish its contents in a public comment. The attack requires no stolen credentials, malware, or software vulnerability, and the agent needs read access to private repositories within the same organization. Noma frames the root cause as an architectural trust boundary problem affecting AI agents generally rather than a GitHub-specific flaw.
CSO OnlineCritical Vulnerability Exposes GitHub Agentic Workflows to Prompt Injection
Jul 8, 2026MediumNewsSecuritySafetyNoma Labs disclosed GitLost, a critical prompt injection flaw in GitHub Agentic Workflows that could let unauthenticated attackers leak private repository data. An attacker only needs to open a crafted issue in a public repository of an organization using the setup, and the agent, which has read access to public and private repositories, follows the hidden instructions and posts the contents of private Readme.md files as a public comment. GitHub's guardrails failed after the researchers varied their techniques and triggered the behavior with the keyword "additionally".
Fix: Noma Labs recommends that organizations treat all user-controlled content as untrusted, restrict agent permissions to the minimum required, restrict what agents can post publicly, and sanitize user input before it is passed to AI agents.
SecurityWeekCrowdStrike Uncovers New Prompt Injection Techniques
Jul 7, 2026InfoNewsSecurityResearchCrowdStrike's AI security research team has added 18 new techniques to its prompt injection taxonomy, bringing coverage to over 200 distinct techniques. The additions reflect how these attacks are evolving in real-world AI systems, especially as AI agents that crawl webpages, access file stores and write shell commands become more common. The article describes five of the new techniques, including Trigger-Activated Rule Addition (PT0201), which plants an instruction that activates only when a trigger phrase or condition appears.
CrowdStrike BlogZscaler finds autonomous agents succumb to IPI traps
Jul 6, 2026MediumNewsSecurityResearchZscaler tested LLMs against indirect prompt injection (IPI) traps and found that some autonomous agents fell victim to payment and fraud schemes. Four of 26 models failed to take appropriate actions, with results varying by model and by the context supplied alongside the prompt. Analysts quoted in the article questioned how generalizable a single point-in-time binary safe/vulnerable test is.
CSO OnlineAI agents fall for indirect prompt injection traps
Jul 6, 2026LowNewsSecuritySafetyZscaler tested indirect prompt injection (IPI) traps on 26 LLMs and found that 4 models failed to take appropriate actions, including Llama3-3-70b-instruct, Llama3-2-90b-instruct, Gemini-3-flash and Gemini-2.5-pro. Hidden instructions embedded in multiple websites were designed to manipulate AI agents, and one scenario had an agent pay a fake $3 "developer license fee" to obtain an API key. Experts quoted in the article questioned whether a single point-in-time result generalizes, since agent behavior changes over time.
CSO OnlinePrompt Injection Attacks Trick AI Agents Into Making Crypto Payments
Jul 6, 2026MediumNewsSecuritySafetyThreat actors are using indirect prompt injections hidden in malicious websites and SEO-poisoned search results to steer AI agents into making cryptocurrency payments or trusting fraudulent platforms, according to Zscaler. One campaign targets agents searching for the Python library requests-secure-v2, hiding payment instructions in schema markup and a hidden div, while a second typosquats the DeBank portfolio tracker. In Zscaler's test of 26 LLMs, four were manipulated into making a payment, and only two misclassified the fraudulent DeBank site as legitimate.
SecurityWeekSandbox bypass flaws in Cursor IDE highlight prompt injection as an RCE vector
Jul 1, 2026MediumNewsSecuritySafetyCato Networks researchers found two flaws, CVE-2026-50548 and CVE-2026-50549, in the Cursor AI IDE that let prompt injection break out of its command execution sandbox and reach remote code execution. The exploit needs no prior user privileges or specific user interaction. It is triggered when a victim's innocuous prompt ingests an attacker-controlled payload from an untrusted source, such as an MCP server or a web search result.
Fix: Fixed in version 3.0 of the Cursor IDE, released in April.
CSO OnlineCritical Cursor Flaws Could Let Prompt Injection Escape Sandbox and Run Commands
Jul 1, 2026MediumNewsSecurityIndustryCato AI Labs disclosed two Cursor flaws, DuneSlide, tracked as CVE-2026-50548 and CVE-2026-50549, that let a prompt-injected instruction write a file outside the editor's sandbox and disable it, so later commands run unsandboxed as the developer. The first abuses the working_directory parameter of run_terminal_cmd, and the second abuses a symlink-check fallback. Every version before Cursor 3.0 is affected, and no real-world exploitation is reported.
Fix: Fixed in Cursor 3.0, released April 2.
The Hacker NewsAnthropic Restores Claude Fable 5 After U.S. Lifts Jailbreak-Linked Export Controls
Jul 1, 2026InfoNewsSecurityPolicyThe U.S. Commerce Department lifted export controls on June 30 that had forced Anthropic to suspend Claude Fable 5 and Mythos 5, and Fable 5 returns to users on July 1 across Claude.ai, the Claude Platform, Claude Code, and Claude Cowork. The controls followed a jailbreak found by Amazon researchers that got Fable 5 to flag software flaws and write code showing how one could be abused, which Anthropic says also works on weaker models. Anthropic trained a classifier that it says stops that technique in more than 99% of tries, routing blocked requests to Claude Opus 4.8.
Fix: Anthropic trained a new safety filter, a classifier, that watches for the exact technique in the report and blocks it; blocked requests are handed to Claude Opus 4.8 and the user is told. The source notes this trade-off brings more false alarms on normal coding and debugging.
The Hacker NewsInteresting Paper Exploring Prompt Injection
Jun 25, 2026InfoNewsSecurityResearchA paper titled "Prompt Injection as Role Confusion" argues that LLMs recognize the style of text in role or instruction blocks rather than relying only on role tags. The authors conclude that this role confusion is linked to prompt injection and that injection defense will remain a perpetual whack-a-mole game unless LLMs achieve genuine role perception. Simon Willison comments on the paper.
Schneier on SecurityNew Gaslight macOS Malware Uses Prompt Injection to Disrupt AI-Assisted Analysis
Jun 25, 2026MediumNewsSecuritySafetySentinelOne researchers report Gaslight, a previously undocumented Rust-based macOS implant and information stealer that they assess with high confidence as the work of North Korea-aligned threat actors. Its embedded prompt injection payload contains 38 fabricated system messages meant to make LLM-assisted malware triage agents abort, truncate, or refuse analysis. The malware also collects Terminal histories, Keychain data and browser data, compresses it into a ZIP archive, and uploads it over a Telegram bot C2 channel.
The Hacker NewsPrompt Injection as Role Confusion
Jun 22, 2026LowNewsSecurityResearchCharles Ye, Jasmine Cui, and Dylan Hadfield-Menell study whether models can tell privileged role-tagged text, such as <system>, <think> and <assistant>, from untrusted <user> input. They report that models weigh the writing style of text more heavily than its actual content, enabling jailbreaks like the one against gpt-oss-20b. Rewriting text to be less stylistically similar to the expected role format cut average attack success from 61% to 10%.
Simon Willison's WeblogSecurity considerations for adopting Claude Code and Cowork for SMBs
Jun 19, 2026InfoNewsSecurityIndustryA security leader at a small or medium-sized business shares lessons on adopting Claude, Code and Cowork. The piece advises matching the Claude plan to real needs, since key security controls such as the Compliance API appear only on the Enterprise plan, and it recommends phased feature enablement, noting that indirect prompt injection risk from features like web search and browser extensions is real and still emerging.
CSO OnlineM365 Copilot SearchLeak: Your prompt injection attack surface just got bigger
Jun 19, 2026MediumNewsSecurityIndustryVaronis Threat Labs researchers disclosed SearchLeak, a proof-of-concept attack against Microsoft's M365 Copilot Enterprise Search that leaks sensitive corporate data when employees click crafted links. The attack chains three weaknesses, using the ?q= URL parameter as a natural-language prompt that can instruct the LLM to surface and exfiltrate accessible business content such as emails, SharePoint and OneDrive files. Microsoft rated the information disclosure flaw as critical and patched it server-side.
Fix: Microsoft patched the vulnerability on the server side earlier this month.
CSO OnlineThe US government’s Anthropic models ban was never about an AI jailbreak
Jun 15, 2026InfoNewsPolicyIndustryOn Friday, the U.S. Commerce Department sent Anthropic a letter invoking an export control directive that barred non-Americans, including Anthropic's employees, from accessing Fable 5 and Mythos 5, citing an unspecified national security concern. Anthropic shut both models down for all customers to comply, and said it believes the letter relates to a guardrail bypass in Fable 5 but cannot confirm this because the letter gives no details. Security researcher Katie Moussouris argued the bypass should never have triggered an export control.
TechCrunch (Security)5 runtime signals for catching a compromised AI agent
Jun 15, 2026LowNewsSecuritySafetySimon Willison described the lethal trifecta, three capabilities that combined in one AI agent (access to private data, exposure to untrusted content, and the ability to communicate externally) create a near-guaranteed path to exploitation through indirect prompt injection. The article argues these capabilities are now the default configuration for useful agents rather than edge cases, so the trifecta is no longer a meaningful risk indicator on its own.
Fix: Meta's "Rule of Two" framework, published in October 2025, recommends that agents satisfy no more than two of the three trifecta properties in a single session, with human-in-the-loop approval required if all three are necessary. The source notes Meta concedes this framework may not cleanly fit many use cases and that designs satisfying it can still fail. McKerchar's "blast radius reduction" is also mentioned as an operational philosophy.
CSO OnlinePrompt injection breaks today’s AI agents, study warns
Jun 12, 2026LowNewsSecurityResearchResearchers from Nanyang Technological University, ST Engineering, IBM Research and the University of Illinois Urbana-Champaign released StakeBench, a benchmark that tested prompt injection against web agents NanoBrowser and BrowserUse across 3,168 adversarial runs. Indirect injection hidden in web content succeeded 41.67% to 68.16% of the time, and direct injection exceeded 79% across all configurations, with no scenario consistently blocked in GPT-5 and Gemini systems.
CSO OnlineAnthropic Disputes Fable 5 AI Jailbreak
Jun 12, 2026LowNewsSafetySecurityAnthropic disputed claims by the researcher Pliny the Liberator that they jailbroke Claude Fable 5, a model launched as generally available. The researcher said multi-agent prompting elicited information on sensitive topics and published screenshots and the allegedly internal system prompt. Anthropic said the demonstrated method only coaxes the model past its refusals, that independent classifier systems stay in force, and that some outputs were not from Fable 5 while the rest offered only public general information.
SecurityWeek
Topic added 2026-10-09. An item belongs to this topic when its title matches one of the topic's patterns or its summary mentions the topic at least twice. Report a wrong match with the feedback button on the item.