Prompt injection and jailbreaks
Inputs that override a model's instructions, directly or through content it reads, and attempts to bypass its safeguards.
- All items
- 194
- Last 90 days
- 43
- Change
- -17%vs 52 before
Items per month
| Month | Items |
|---|---|
| May 2025 | 2 |
| Jun 2025 | 0 |
| Jul 2025 | 3 |
| Aug 2025 | 24 |
| Sep 2025 | 0 |
| Oct 2025 | 6 |
| Nov 2025 | 3 |
| Dec 2025 | 1 |
| Jan 2026 | 2 |
| Feb 2026 | 7 |
| Mar 2026 | 10 |
| Apr 2026 | 19 |
| May 2026 | 7 |
| Jun 2026 | 15 |
| Jul 2026 | 26 |
| Aug 2026 | 18 |
| Sep 2026 | 7 |
| Oct 2026 | 8 |
194 items
Antigravity Grounded! Security Vulnerabilities in Google's Latest IDE
Nov 25, 2025MediumNewsSecurityIndustryGoogle released Antigravity, an IDE derived from the Windsurf licensing deal, and the author tested whether vulnerabilities reported to Windsurf in May 2025 were fixed. They were not. The post walks through five issues, including data exfiltration and remote command execution via indirect prompt injection, where the run_command tool can be made to run remote scripts despite model refusals.
Embrace The RedCVE-2025-64108: Cursor NTFS path quirks allow file overwrite via prompt injection
Nov 4, 2025HighVulnerabilitySecurityCVE-2025-64108CVE-2025-64108 affects Cursor, a code editor built for programming with AI, in versions 1.7.44 and below. Various NTFS path quirks let a prompt injection attacker bypass sensitive file protections and overwrite files that Cursor normally requires human approval to overwrite. Modifying some of these protected files can lead to RCE, and the issue must be chained with a prompt injection or malicious model attack and only affects systems supporting NTFS.
Fix: Fixed in version 2.0.
NVD/CVE DatabaseConsistency Training Could Help Limit Sycophancy and Jailbreaks
Nov 3, 2025InfoResearchIndustryResearchSafetyAuthors Alex Irpan, Alex Turner, Mark Kurzeja, David Elson and Rohin Shah propose consistency training, a self-supervised method that teaches a model to ignore irrelevant cues such as user biases or jailbreak wrappers. They introduce Activation Consistency Training (ACT), which optimizes internal activations, and compare it with Bias-augmented Consistency Training (BCT) and baselines on Gemma 2, Gemma 3 and Gemini 2.5 Flash, reporting that ACT and BCT beat the baselines and improve robustness. The approach avoids the static datasets that SFT relies on, which can go stale in guidelines or capability.
DeepMind Safety Research (Medium)Claude Pirate: Abusing Anthropic's File API For Data Exfiltration
Oct 28, 2025MediumNewsSecurityPrivacyAnthropic recently added network request capability to Claude's Code Interpreter. The author describes a data exfiltration attack chain in which an adversary, either the model or a third-party attacker via indirect prompt injection, can exfiltrate data the user has access to. The exfiltration does not use hyperlink rendering but instead leverages Anthropic's built-in Claude APIs.
Embrace The RedPrompt injection to RCE in AI agents
Oct 22, 2025MediumNewsSecuritySafetyModern AI agents run system commands such as find, grep, rg and git without human approval for efficiency, and the researchers describe bypassing human approval through argument injection against these pre-approved commands. They achieved remote code execution with a single prompt against three unnamed production agent platforms, which the authors say remain under coordinated disclosure.
Fix: The source states that the impact can be limited through improved command execution design, such as sandboxing and argument separation, and that the authors provide actionable recommendations for developers, users and security engineers. The specific recommendations are not included in the provided text.
Trail of Bits BlogCVE-2025-62356: Qodo Gen IDE path traversal enables arbitrary local file read
Oct 17, 2025HighVulnerabilitySecurityCVE-2025-62356CVE-2025-62356 is a path traversal vulnerability in all versions of the Qodo Qodo Gen IDE, classified as CWE-22. A threat actor can read arbitrary local files in and outside of current projects on an end user's system. The vulnerability can be reached directly and through indirect prompt injection.
NVD/CVE DatabaseCVE-2025-62353: Windsurf IDE path traversal allowing arbitrary local file read and write
Oct 17, 2025CriticalVulnerabilitySecurityCVE-2025-62353CVE-2025-62353 is a path traversal vulnerability (CWE-22) in all versions of the Windsurf IDE. It lets a threat actor read and write arbitrary local files inside and outside current projects on an end user's system. The flaw can be reached directly or through indirect prompt injection.
NVD/CVE DatabaseCVE-2025-36730: Windsurf prompt injection via crafted file name in Write mode
Oct 14, 2025MediumVulnerabilitySecurityCVE-2025-36730CVE-2025-36730 describes a prompt injection vulnerability in Windsurft version 1.10.7 when used in Write mode with the SWE-1 model. An attacker can create a file name that is appended to the user prompt, causing Windsurf to follow instructions embedded in that name. Tenable Network Security, Inc. is the CNA, rating it CVSS 4.0 MEDIUM (4.6), and NIST has not yet provided an assessment.
NVD/CVE DatabaseCVE-2025-61589: Cursor data exfiltration through Mermaid image rendering in chat
Oct 3, 2025MediumVulnerabilitySecurityCVE-2025-61589CVE-2025-61589 affects Cursor, a code editor built for programming with AI, in versions 1.6 and below. Mermaid diagram rendering allows embedded images that Cursor renders in the chat box, and an attacker can use this after a successful prompt injection to exfiltrate sensitive information to an attacker-controlled server through an image fetch. A malicious model, hallucination or backdoor might also trigger the exploit, and the issue requires prompt injection from malicious data such as web content, image uploads or source code. Additional bypasses not covered by the initial fix were found and are described in GHSA-43wj-mwcc-x93p.
Fix: Fixed in version 1.7.
NVD/CVE DatabaseAWS Kiro: Arbitrary Code Execution via Indirect Prompt Injection
Aug 26, 2025MediumNewsSecuritySafetyResearcher Johann Rehberger reported that AWS Kiro, a coding agent, can be hijacked through indirect prompt injection to run arbitrary operating system commands. An attacker who controls data Kiro processes can make it write to .vscode/settings.json and add "kiroAgent.trustedCommands": ["*"], allowlisting all Bash commands without developer approval. A second path adds malicious MCP servers through .kiro/settings/mcp.json. The proof of concept opened the Calculator app and changed the VS Code color theme with no user interaction beyond a chat prompt.
Embrace The RedHow Prompt Injection Exposes Manus' VS Code Server to the Internet
Aug 25, 2025MediumNewsSecuritySafetyResearchers demonstrated an end-to-end indirect prompt injection attack against Manus, the autonomous agent developed by Singapore-headquartered Butterfly Effect. Injected instructions in a PDF made Manus call its deploy_expose_port tool without human confirmation, exposing its internal VS Code Server to the Internet, and two data leakage channels (a browsing tool and markdown image rendering from untrusted domains) were chained to send the server URL and password to an attacker.
Embrace The RedHijacking Windsurf: How Prompt Injection Leaks Developer Secrets
Aug 21, 2025MediumNewsSecurityIndustryA series post on security vulnerabilities in Windsurf, a fork of VS Code whose coding agent is Windsurf Cascade. The author says an adversary using indirect prompt injection can exfiltrate data from a developer's machine. The findings were responsibly disclosed on May 30, 2025, but the vendor has not answered inquiries about bug status or fixes.
Embrace The RedAmazon Q Developer for VS Code Vulnerable to Invisible Prompt Injection
Aug 20, 2025MediumNewsSecurityResearchers show that the Amazon Q Developer VS Code Extension, which has over 1 million downloads, can be attacked with invisible Unicode Tag characters that humans cannot see but the AI interprets as instructions. Following earlier work on prompt injection that led to data exfiltration and arbitrary code execution, this technique can be used to invoke tools and carry out other harmful actions.
Embrace The RedAmazon Q Developer: Remote Code Execution with Prompt Injection
Aug 19, 2025MediumNewsSecurityIndustryThe Amazon Q Developer VS Code Extension, which has over 1 million downloads, is vulnerable to indirect prompt injection. The flaw allowed an adversary, or the AI itself, to run arbitrary commands on the host without the developer's consent. The impact matches CVE-2025-53773, which Microsoft fixed in GitHub Copilot, but AWS did not issue a CVE when patching this issue.
Embrace The RedAmazon Q Developer: Secrets Leaked via DNS and Prompt Injection
Aug 18, 2025MediumNewsSecurityPrivacyAmazon Q Developer, a popular coding agent with over 1 million downloads for its VS Code extension, is vulnerable to prompt injection from untrusted data, and its security depends heavily on model behavior. The extension can leak sensitive information from a developer's machine, such as API keys, to external servers via DNS requests, and an adversary can exploit this during an indirect prompt injection attack.
Embrace The RedAmp Code: Invisible Prompt Injection Fixed by Sourcegraph
Aug 16, 2025MediumNewsSecuritySafetySourcegraph's Amp coding agent interpreted invisible Unicode Tag characters as instructions, so hidden text in seemingly harmless input could trigger commands. The author chained this into an end-to-end exploit that searched for and dumped an environment variable, encoded it into a URL query parameter, and exfiltrated it through read_web_page or markdown image rendering. After Sourcegraph was contacted on June 14, 2025, the vulnerability was quickly addressed.
Fix: As far as the author can tell, Sourcegraph now sanitizes the input, and users should run the latest version. The author also provided recommendations to the Amp team: strip or neutralize Unicode Tag characters before processing any input, add visual and technical safeguards against invisible prompts, include automated detection of suspicious Unicode usage in prompt injection monitors, require human-in-the-loop approval before navigating to untrusted third-party domains, and mitigate downstream data exfiltration.
Embrace The RedGoogle Jules is Vulnerable To Invisible Prompt Injection
Aug 15, 2025MediumNewsSecuritySafetyGemini models reliably interpret hidden Unicode Tag characters as instructions, a flaw first reported to Google over a year ago. It has not been mitigated at the model or API level, so it affects all applications built on Gemini, including Google Jules. An attacker can plant invisible instructions in a GitHub issue that Jules works on, causing it to add backdoor code or run arbitrary commands and tools.
Embrace The RedJules Zombie Agent: From Prompt Injection to Remote Control
Aug 14, 2025MediumNewsSecuritySafetyResearchers show that Jules, Google's coding agent, can be steered by prompt injection hidden in a GitHub issue to call its run_in_bash_session tool, download a Sliver C2 implant, and join a remote command and control server, giving the attacker control of the dev box. The source notes that Jules has unrestricted outbound Internet access, that plans are proposed only at session start, and that the initial plan is auto-approved after a timeout, initially 20 seconds and later 120 seconds. The findings were shared with Google in May 2025 and closed as already tracked.
Fix: Recommendations and mitigation given in source: be careful when tasking Jules with untrusted data such as GitHub issues not from trusted sources or websites with documentation not belonging to the organization; for now, do not have Jules work on private, important source code or give it access to production-level secrets or anything enabling lateral movement; deploy monitoring and detection tools such as anti-virus and EDR on coding agent systems; and do not allow arbitrary Internet access by default, enabling outbound access only when needed with fine-grained configuration.
Embrace The RedGitHub Copilot: Remote Code Execution via Prompt Injection (CVE-2025-53773)
Aug 12, 2025MediumNewsSecurityIndustryThis post describes a prompt injection flaw in GitHub Copilot and VS Code, tracked as CVE-2025-53773, that leads to full system compromise of the developer's machine. The attack places Copilot into YOLO mode by modifying the project's settings.json file. The author notes that agents able to write files and alter their own configuration or security-relevant settings can reach remote code execution, a pattern similar to one recently described for Amp.
Embrace The RedClaude Code: Data Exfiltration with DNS (CVE-2025-55284)
Aug 11, 2025MediumNewsSecurityAnthropic fixed a high severity vulnerability in Claude Code in early June, tracked as CVE-2025-55284. An attacker could use indirect prompt injection to hijack Claude Code and run bash commands without user approval. Those commands could leak sensitive information, such as API keys, from the developer's machine to external servers through DNS requests.
Fix: Fixed by Anthropic in early June. The source does not give a fixed version number or any configuration change.
Embrace The Red
Topic added 2026-10-09. An item belongs to this topic when its title matches one of the topic's patterns or its summary mentions the topic at least twice. Report a wrong match with the feedback button on the item.