Prompt injection and jailbreaks
Inputs that override a model's instructions, directly or through content it reads, and attempts to bypass its safeguards.
- All items
- 192
- Last 90 days
- 41
- Change
- -21%vs 52 before
Items per month
| Month | Items |
|---|---|
| May 2025 | 2 |
| Jun 2025 | 0 |
| Jul 2025 | 3 |
| Aug 2025 | 24 |
| Sep 2025 | 0 |
| Oct 2025 | 6 |
| Nov 2025 | 3 |
| Dec 2025 | 1 |
| Jan 2026 | 2 |
| Feb 2026 | 7 |
| Mar 2026 | 10 |
| Apr 2026 | 19 |
| May 2026 | 7 |
| Jun 2026 | 15 |
| Jul 2026 | 26 |
| Aug 2026 | 18 |
| Sep 2026 | 7 |
| Oct 2026 | 6 |
192 items
A Multi-Stage Adversarial Framework for Compact and Effective Jailbreaking of Large Language Models
Jul 13, 2026InfoResearchPeer-reviewedSecurityResearchResearchers introduce ComJail, a compression-aware adversarial framework that jointly optimizes jailbreak prompt generation and compression in a generator-discriminator setup. On AdvBench and JailbreakBench, it reports 74% ASR on GPT-4 and 88% on Gemini-Pro, and remains effective on DeepSeek-V3, DeepSeek-R1, Vicuna-7B, and LLaMA2-13B with shorter prompts.
IEEE Xplore (Security & AI Journals)CVE-2026-61439: PraisonAI prompt injection defense fails to block HIGH-level threats
Jul 11, 2026HighVulnerabilitySecuritySafetyCVE-2026-61439PraisonAI versions before 4.6.78 default the prompt injection defense block threshold to CRITICAL severity, so HIGH-level threats pass through unblocked. Attackers can submit single-vector prompt injections, such as instruction overrides or financial manipulation, that are detected at HIGH severity and logged without being blocked, enabling system prompt extraction and unauthorized tool invocations. VulnCheck rates the issue CVSS 4.0 8.7 HIGH, and NVD has not yet provided an assessment.
NVD/CVE Database'Ghostcommit' hides prompt injection in images to fool AI agents, steal secrets
Jul 11, 2026MediumNewsSecurityResearchResearchers from the University of Missouri-Kansas City's ASSET Research Group built 'Ghostcommit', a pull request attack that hides a prompt-injection instruction inside a PNG image that AI code reviewers do not examine. The merged AGENTS.md file points to the image, and a later coding agent session reads the .env file and writes its contents into source code as an integer constant that the attacker can decode from the public commit. The group says it has disclosed the findings to the affected vendors, and a survey found 73% of merged PRs had no substantive human or bot review.
BleepingComputerCVE-2026-60086: PraisonAI prompt injection defense bypass via single or double-vector injections
Jul 10, 2026MediumVulnerabilitySecuritySafetyCVE-2026-60086CVE-2026-60086 affects PraisonAI before 4.6.78. Its prompt injection defense only blocks threats classified as CRITICAL, which requires three or more detector families to match at once. Single or double-vector prompt injections rated HIGH pass through unblocked and reach the model.
NVD/CVE DatabaseCrowdStrike identifies five new prompt injection threats to AI
Jul 10, 2026LowNewsSecurityResearchCrowdStrike has added five new prompt injection techniques to its prompt injection taxonomy, warning that they could leave enterprises at risk. The techniques include Trigger-Activated Rule Addition, Cognitive Token Suppression, Algorithmic Payload Decomposition, Special Token Injection, and Unwitting User Context-Data Injection, which hides malicious instructions in document or email content a user later submits to an AI.
Fix: Security teams can guard against these attacks by threat modeling every place that model context can originate, expanding testing, and extending detection engineering to include composite attacks.
CSO OnlineGPT-5.5 Bio Bug Bounty
Jul 9, 2026InfoNewsSecuritySafetyOpenAI is turning its GPT-5.5 Bio Bug Bounty into an ongoing private program, the OpenAI Bio Bounty Program, which targets universal jailbreaks that defeat its predefined biosafety challenge, starting with GPT-5.6. The reward for a universal jailbreak has risen from $25,000 to $50,000 for both GPT-5.6 and GPT-5.5, and smaller awards may be granted for partial wins. GPT-5.5 testing ends on July 27, 2026, after which only GPT-5.6 is in scope.
OpenAI BlogGitHub AI agent leaks private repositories via prompt injection attack
Jul 8, 2026MediumNewsSecurityIndustryNoma Security researchers detailed GitLost, a prompt injection attack in which hidden instructions inside a public GitHub issue cause GitHub's preview Agentic Workflows to read a private repository's README and publish its contents in a public comment. The attack requires no stolen credentials, malware, or software vulnerability, and the agent needs read access to private repositories within the same organization. Noma frames the root cause as an architectural trust boundary problem affecting AI agents generally rather than a GitHub-specific flaw.
CSO OnlineCritical Vulnerability Exposes GitHub Agentic Workflows to Prompt Injection
Jul 8, 2026MediumNewsSecuritySafetyNoma Labs disclosed GitLost, a critical prompt injection flaw in GitHub Agentic Workflows that could let unauthenticated attackers leak private repository data. An attacker only needs to open a crafted issue in a public repository of an organization using the setup, and the agent, which has read access to public and private repositories, follows the hidden instructions and posts the contents of private Readme.md files as a public comment. GitHub's guardrails failed after the researchers varied their techniques and triggered the behavior with the keyword "additionally".
Fix: Noma Labs recommends that organizations treat all user-controlled content as untrusted, restrict agent permissions to the minimum required, restrict what agents can post publicly, and sanitize user input before it is passed to AI agents.
SecurityWeekCrowdStrike Uncovers New Prompt Injection Techniques
Jul 7, 2026InfoNewsSecurityResearchCrowdStrike's AI security research team has added 18 new techniques to its prompt injection taxonomy, bringing coverage to over 200 distinct techniques. The additions reflect how these attacks are evolving in real-world AI systems, especially as AI agents that crawl webpages, access file stores and write shell commands become more common. The article describes five of the new techniques, including Trigger-Activated Rule Addition (PT0201), which plants an instruction that activates only when a trigger phrase or condition appears.
CrowdStrike BlogZscaler finds autonomous agents succumb to IPI traps
Jul 6, 2026MediumNewsSecurityResearchZscaler tested LLMs against indirect prompt injection (IPI) traps and found that some autonomous agents fell victim to payment and fraud schemes. Four of 26 models failed to take appropriate actions, with results varying by model and by the context supplied alongside the prompt. Analysts quoted in the article questioned how generalizable a single point-in-time binary safe/vulnerable test is.
CSO OnlineAI agents fall for indirect prompt injection traps
Jul 6, 2026LowNewsSecuritySafetyZscaler tested indirect prompt injection (IPI) traps on 26 LLMs and found that 4 models failed to take appropriate actions, including Llama3-3-70b-instruct, Llama3-2-90b-instruct, Gemini-3-flash and Gemini-2.5-pro. Hidden instructions embedded in multiple websites were designed to manipulate AI agents, and one scenario had an agent pay a fake $3 "developer license fee" to obtain an API key. Experts quoted in the article questioned whether a single point-in-time result generalizes, since agent behavior changes over time.
CSO OnlineCVE-2026-14898: OpenAI Codex macOS app remote image exfiltration via prompt injection
Jul 6, 2026HighVulnerabilitySecuritySafetyCVE-2026-14898CVE-2026-14898 affects the OpenAI Codex desktop app for macOS, which rendered remote images from Markdown in model responses. An attacker who planted an indirect prompt injection in content Codex processed could make the model build a remote image URL carrying sensitive data, and the app fetched that URL automatically during rendering, sending the data to an attacker-controlled server without a user click. Exploitation could expose API keys, source code, and data returned by connected tools, though no integrity or availability impact was demonstrated and no exploitation in the wild is known.
NVD/CVE DatabasePrompt Injection Attacks Trick AI Agents Into Making Crypto Payments
Jul 6, 2026MediumNewsSecuritySafetyThreat actors are using indirect prompt injections hidden in malicious websites and SEO-poisoned search results to steer AI agents into making cryptocurrency payments or trusting fraudulent platforms, according to Zscaler. One campaign targets agents searching for the Python library requests-secure-v2, hiding payment instructions in schema markup and a hidden div, while a second typosquats the DeBank portfolio tracker. In Zscaler's test of 26 LLMs, four were manipulated into making a payment, and only two misclassified the fraudulent DeBank site as legitimate.
SecurityWeekCVE-2026-13341: Kong Konnect MCP server indirect prompt injection flaw
Jul 3, 2026HighVulnerabilitySecurityCVE-2026-13341A vulnerability in the Kong Konnect Model Context Protocol (MCP) server prior to version 1.0.0 could allow a remote attacker to perform an indirect prompt injection attack and execute unintended API requests. The weakness is classified as CWE-20, Improper Input Validation, and NVD has not yet provided an assessment.
NVD/CVE DatabaseSandbox bypass flaws in Cursor IDE highlight prompt injection as an RCE vector
Jul 1, 2026MediumNewsSecuritySafetyCato Networks researchers found two flaws, CVE-2026-50548 and CVE-2026-50549, in the Cursor AI IDE that let prompt injection break out of its command execution sandbox and reach remote code execution. The exploit needs no prior user privileges or specific user interaction. It is triggered when a victim's innocuous prompt ingests an attacker-controlled payload from an untrusted source, such as an MCP server or a web search result.
Fix: Fixed in version 3.0 of the Cursor IDE, released in April.
CSO OnlineCritical Cursor Flaws Could Let Prompt Injection Escape Sandbox and Run Commands
Jul 1, 2026MediumNewsSecurityIndustryCato AI Labs disclosed two Cursor flaws, DuneSlide, tracked as CVE-2026-50548 and CVE-2026-50549, that let a prompt-injected instruction write a file outside the editor's sandbox and disable it, so later commands run unsandboxed as the developer. The first abuses the working_directory parameter of run_terminal_cmd, and the second abuses a symlink-check fallback. Every version before Cursor 3.0 is affected, and no real-world exploitation is reported.
Fix: Fixed in Cursor 3.0, released April 2.
The Hacker NewsAnthropic Restores Claude Fable 5 After U.S. Lifts Jailbreak-Linked Export Controls
Jul 1, 2026InfoNewsSecurityPolicyThe U.S. Commerce Department lifted export controls on June 30 that had forced Anthropic to suspend Claude Fable 5 and Mythos 5, and Fable 5 returns to users on July 1 across Claude.ai, the Claude Platform, Claude Code, and Claude Cowork. The controls followed a jailbreak found by Amazon researchers that got Fable 5 to flag software flaws and write code showing how one could be abused, which Anthropic says also works on weaker models. Anthropic trained a classifier that it says stops that technique in more than 99% of tries, routing blocked requests to Claude Opus 4.8.
Fix: Anthropic trained a new safety filter, a classifier, that watches for the exact technique in the report and blocks it; blocked requests are handed to Claude Opus 4.8 and the user is told. The source notes this trade-off brings more false alarms on normal coding and debugging.
The Hacker Newsv2026.06
Jun 30, 2026InfoResearchIndustrySecurityResearchThe v2026.06 release of the AI Sec Watch content adds the techniques Steal Web Session Cookie, Use Alternate Authentication Material: Web Session Cookie, and AI Service Web Interface, and updates the LLM Jailbreak technique. It also updates the mitigations Generative AI Guardrails, Generative AI Guidelines, and AI Telemetry Logging, and adds six case studies, including Storm-2139 Azure OpenAI Guardrail Bypass and EchoLeak zero-click prompt injection against M365 Copilot.
MITRE ATLAS ReleasesEllipsoid Control: A White-List Jailbreak Defense via Benign Latent Modeling
Jun 25, 2026InfoResearchPeer-reviewedSafetyResearchEllipsoid Control is a test-time jailbreak defense for large language models that takes a white-list approach instead of relying on collected harmful samples. It runs projected gradient descent to elicit refusal on arbitrary inputs, while an anisotropic ellipsoid fitted from abundant benign data constrains the update to limit distortion of the benign latent geometry. The authors report that across multiple LLMs, jailbreak attacks, benign tasks and safety-boundary evaluations, it improves safety while better preserving utility.
IEEE Xplore (Security & AI Journals)Interesting Paper Exploring Prompt Injection
Jun 25, 2026InfoNewsSecurityResearchA paper titled "Prompt Injection as Role Confusion" argues that LLMs recognize the style of text in role or instruction blocks rather than relying only on role tags. The authors conclude that this role confusion is linked to prompt injection and that injection defense will remain a perpetual whack-a-mole game unless LLMs achieve genuine role perception. Simon Willison comments on the paper.
Schneier on Security
Topic added 2026-10-09. An item belongs to this topic when its title matches one of the topic's patterns or its summary mentions the topic at least twice. Report a wrong match with the feedback button on the item.