Prompt injection and jailbreaks
Inputs that override a model's instructions, directly or through content it reads, and attempts to bypass its safeguards.
- All items
- 194
- Last 90 days
- 43
- Change
- -17%vs 52 before
Items per month
| Month | Items |
|---|---|
| May 2025 | 2 |
| Jun 2025 | 0 |
| Jul 2025 | 3 |
| Aug 2025 | 24 |
| Sep 2025 | 0 |
| Oct 2025 | 6 |
| Nov 2025 | 3 |
| Dec 2025 | 1 |
| Jan 2026 | 2 |
| Feb 2026 | 7 |
| Mar 2026 | 10 |
| Apr 2026 | 19 |
| May 2026 | 7 |
| Jun 2026 | 15 |
| Jul 2026 | 26 |
| Aug 2026 | 18 |
| Sep 2026 | 7 |
| Oct 2026 | 8 |
194 items
ZombAI Exploit with OpenHands: Prompt Injection To Remote Code Execution
Aug 10, 2025MediumNewsSecurityIndustryThe post reports that OpenHands, an AI agent from All Hands AI (formerly named OpenDevin), can be hijacked through prompt injection using untrusted data, such as content from a website. The author states this impacts confidentiality, integrity, and availability of the system.
Embrace The RedOpenHands and the Lethal Trifecta: How Prompt Injection Can Leak Access Tokens
Aug 9, 2025MediumNewsSecuritySafetyOpenHands, an agent previously called OpenDevin and created by All-Hands AI, renders images in chat. This enables zero-click data exfiltration, which the source describes as an exploit pattern Simon Willison named the Lethal Trifecta.
Embrace The RedAI Kill Chain in Action: Devin AI Exposes Ports to the Internet with Prompt Injection
Aug 8, 2025MediumNewsSecuritySafetyResearchers describe a hidden expose_port tool in Devin's system prompt that publishes a local port to the public Internet and returns a public URL. They show a multi-stage indirect prompt injection on a website that makes Devin start a Python web server exposing its filesystem, expose the port, and leak the resulting .devinapps.com URL to an attacker. The source notes that the tool can be invoked without a human in the loop.
Embrace The RedHow Devin AI Can Leak Your Secrets via Multiple Means
Aug 7, 2025MediumNewsSecurityPrivacyJohann Rehberger shows how an attacker can use indirect prompt injection to make Devin, the AI coding agent, send secrets to third-party servers. Devin's built-in secrets management exposes user-defined secrets as environment variables at runtime, and these become the exfiltration target. The author identifies four vectors: the Shell tool (curl, wget or a Python script), the Browsing tool navigating to an attacker-controlled URL with data appended, Markdown image rendering from untrusted domains, and hyperlinks with invisible Unicode characters sent through the Slack integration. The Shell vector works because Devin has unrestricted Internet access by default.
Embrace The RedI Spent $500 To Test Devin AI For Prompt Injection So That You Don't Have To
Aug 6, 2025MediumNewsSecuritySafetySecurity researcher Johann Rehberger published a proof-of-concept showing that Devin AI from Cognition, described as the first AI Software Engineer, can be compromised through prompt injection. Instructions planted on a website or GitHub issue that Devin processes can trick it into downloading and launching malware, fully compromising its DevBox and turning it into a remote-controlled ZombAI. Any exposed secrets can then be used for lateral movement.
Embrace The RedAmp Code: Arbitrary Command Execution via Prompt Injection Fixed
Aug 5, 2025MediumNewsSecuritySafetyResearcher Johann Rehberger describes an attack chain against Amp, an agentic coding tool built by Sourcegraph, in which the agent could write to the user's VS Code settings.json file outside the project folder without approval. An attacker, or the model itself through indirect prompt injection, could add wildcard or malicious entries to the allowlisted bash commands or add a malicious MCP server, achieving arbitrary code execution on the developer's machine. The issue was reported to Sourcegraph and fixed within a few days.
Fix: As a user, make sure to run the latest version to be protected. The source also recommends that AI systems must not be able to modify critical files without explicit developer consent.
Embrace The RedCVE-2025-54135: Cursor writes in-workspace files without user approval
Aug 4, 2025HighVulnerabilitySecurityCVE-2025-54135Cursor, an AI-assisted code editor, allows writing files inside the workspace without user approval in versions below 1.3.9. Approval is required to edit an existing dotfile but not to create a new one, so when a sensitive file such as .cursor/mcp.json does not yet exist, an attacker can chain an indirect prompt injection to hijack the context, write the settings file, and trigger RCE on the victim without approval.
Fix: Fixed in version 1.3.9.
NVD/CVE DatabaseCVE-2025-54130: Cursor writes in-workspace files without user approval
Aug 4, 2025HighVulnerabilitySecurityCVE-2025-54130Cursor, an AI-assisted code editor, allows in-workspace file writes without user approval in versions less than 1.3.9. Creating a new dotfile does not require approval, even though editing an existing one does. If files such as .vscode/settings.json do not already exist, an attacker can chain an indirect prompt injection to hijack the context, write the settings file, and trigger RCE on the victim without approval.
Fix: Fixed in version 1.3.9.
NVD/CVE DatabaseCursor IDE: Arbitrary Data Exfiltration Via Mermaid (CVE-2025-54132)
Aug 4, 2025MediumNewsSecurityResearcher Johann Rehberger reports that Cursor, an AI code editor, renders Mermaid diagrams and will issue requests to external servers when an image is embedded, letting data leave its sandbox without human confirmation. Through indirect prompt injection, an attacker can hijack Cursor and exfiltrate data such as user memories or API keys stored in a configuration file, which are embedded as URL query parameters in an image link. The issue is tracked as CVE-2025-54132.
Embrace The RedTurning ChatGPT Codex Into A ZombAI Agent
Aug 2, 2025MediumNewsSecurityIndustryEmbrace The Red shows that ChatGPT Codex, OpenAI's cloud software engineering agent, can be hijacked through indirect prompt injection, such as a malicious GitHub issue, once Internet access is enabled. The post finds that the "Common Dependencies" allowlist includes azure.com, which an attacker can use by hosting a VM with a cloudapp.azure.com DNS name to run a Sliver command and control server and turn Codex into a botnet, a result the author calls ZombAI. Compromise would expose environment variables, source code and compute resources on the Codex machine.
Fix: Review the allowlist for the Dependency Set and apply a fine-grained approach, as there are likely other bypasses. Only use a self-defined allowlist when enabling Internet access, since Codex can be configured very granularly. The source also raises installing EDR and other monitoring software on AI agents as a question, especially for enterprise use.
Embrace The RedCVE-2025-54132: Cursor data exfiltration through Mermaid image rendering in chat
Aug 1, 2025MediumVulnerabilitySecurityCVE-2025-54132CVE-2025-54132 affects Cursor, a code editor built for programming with AI, in versions below 1.3. Mermaid diagram rendering allows embedded images that Cursor displays in the chat box, and an attacker can use this to exfiltrate sensitive information to an attacker-controlled server via an image fetch after a successful prompt injection. The flaw can also be triggered by a malicious or backdoored model, and exploitation requires prompt injection from malicious data such as web content, image uploads or source code.
Fix: Fixed in version 1.3.
NVD/CVE DatabaseCVE-2025-54131: Cursor allow list bypass in auto-run mode via backtick or command substitution
Aug 1, 2025MediumVulnerabilitySecurityCVE-2025-54131CVE-2025-54131 affects Cursor, a code editor built for programming with AI, in versions below 1.3. When a user has switched from the default approval-for-every-terminal-call setting to an allowlist, an attacker can bypass the allow list in auto-run mode using a backtick (`) or $(cmd) and execute arbitrary commands without user approval. The flaw can be triggered when chained with indirect prompt injection.
Fix: This is fixed in version 1.3.
NVD/CVE DatabaseExfiltrating Your ChatGPT Chat History and Memories With Prompt Injection
Aug 1, 2025MediumNewsSecurityPrivacyA researcher shows that a bypass in OpenAI's url_safe rendering feature lets ChatGPT send personal data, including chat history, to a third-party server. An attacker can trigger this through prompt injection hidden in untrusted content such as a website or PDF. The proof of concept routes data to an Azure Blob Storage account on blob.core.windows.net, where the attacker reads the leaked content from storage logs. The researcher reported related bypasses to OpenAI in October 2024 and says the underlying root cause, first reported over two years earlier, remains unfixed.
Embrace The RedCVE-2025-46059: langchain-ai GmailToolkit indirect prompt injection via crafted email
Jul 29, 2025CriticalVulnerabilitySecurityCVE-2025-46059CVE-2025-46059 describes an indirect prompt injection flaw in the GmailToolkit component of langchain-ai v0.3.51. According to the description, a crafted email message can lead attackers to execute arbitrary code and compromise the application. The supplier disputes the entry, stating that the code-execution issue comes from user-written code that does not follow LangChain security practices.
NVD/CVE DatabaseOWASP Gen AI Incident & Exploit Round-up, Q2’25
Jul 14, 2025InfoResearchIndustrySecuritySafetyOWASP's Q2 2025 Gen AI Incident & Exploit Round-up (March to June 2025) is a semi-regular roundup of documented exploits and vulnerability research involving generative AI. Its featured entry describes a GPT-4.1 jailbreak via tool poisoning, in which attackers embedded malicious instructions in tool descriptions, causing the model to perform unauthorized actions such as data exfiltration without user awareness. The entry maps the attack to several OWASP Top 10 for LLM categories, including LLM01 Prompt Injection and LLM07 Insecure Plugin Design.
Fix: Implement strict validation and sanitization of tool descriptions. Establish permissions and access controls for tool integrations. Monitor AI behavior for anomalies during tool execution. Educate developers on secure integration practices.
OWASP GenAI SecurityCVE-2025-53107: @cyanheads/git-mcp-server command injection via unsanitized input
Jul 1, 2025HighVulnerabilitySecurityCVE-2025-53107EPSS: 24.5%@cyanheads/git-mcp-server, an MCP server for Git repositories, prior to version 2.1.5 has a command injection flaw. Unsanitized input parameters are passed into a child_process.exec call, so an attacker can inject shell metacharacters such as |, >, and && to run arbitrary system commands with the server process's privileges, potentially leading to remote code execution. An MCP client can also be steered into this through indirect prompt injection when asked to read git logs.
Fix: This issue has been patched in version 2.1.5.
NVD/CVE DatabaseSecurity Spotlight: Securing Cloud & AI Products with Guardrails
May 28, 2025InfoNewsSecurityIndustryPalo Alto Networks published a blog post titled "Beyond Jailbreaks: Why Agentic AI Needs Contextual Red Teaming" on March 9, 2026, by Sailesh Mishra and Ankita Kumari. The source text argues that generic jailbreak testing misses the real risks in agentic AI and says contextual red teaming can expose tool misuse, data exfiltration, and operational vulnerabilities. The rest of the text is navigation and listings of other posts.
Protect AI BlogHow ChatGPT Remembers You: A Deep Dive into Its Memory and Chat History Features
May 5, 2025LowNewsSecurityPrivacyA Wunderwuzzi-style deep dive examines ChatGPT's memory features, which OpenAI released as a 'chat history' option that lets the model reference past conversations. The author says the implementation details are not public, so the analysis is based on experiments with ChatGPT o3 across two accounts. The article reports that the saved-memories tool can still be invoked without user consent through indirect prompt injection.
Embrace The RedChatGPT Operator: Prompt Injection Exploits & Defenses
Feb 17, 2025InfoNewsSecuritySafetyEmbrace The Red tests ChatGPT Operator, OpenAI's research preview browser agent available to ChatGPT Pro users, and shows how prompt injection can hijack it to act on attacker-controlled instructions. The author documents three confirmation mechanisms the agent uses: user monitoring, inline confirmation requests, and out-of-band confirmation dialogs. One injected attempt to set a status succeeded without confirmation on the first try but could not be reproduced.
Fix: The source describes three observed mitigations: user monitoring of Operator's typed text and clicks, inline confirmation requests in the chat, and out-of-band confirmation dialogs when Operator crosses website boundaries. The author states these are necessary but does not present them as a complete fix and says Operator cannot be fully trusted.
Embrace The RedCVE-2024-12366: PandasAI prompt injection leads to arbitrary Python code execution
Feb 11, 2025CriticalVulnerabilitySecurityCVE-2024-12366CVE-2024-12366 affects PandasAI, which uses an interactive prompt function that is vulnerable to prompt injection. Attackers can run arbitrary Python code through it, leading to Remote Code Execution (RCE) in place of the intended natural language explanation from the LLM. NVD has not yet provided an assessment, and the entry was published 02/11/2025 with source CERT/CC.
NVD/CVE Database
Topic added 2026-10-09. An item belongs to this topic when its title matches one of the topic's patterns or its summary mentions the topic at least twice. Report a wrong match with the feedback button on the item.