Prompt injection and jailbreaks
Inputs that override a model's instructions, directly or through content it reads, and attempts to bypass its safeguards.
- All items
- 194
- Last 90 days
- 43
- Change
- -17%vs 52 before
Items per month
| Month | Items |
|---|---|
| May 2025 | 2 |
| Jun 2025 | 0 |
| Jul 2025 | 3 |
| Aug 2025 | 24 |
| Sep 2025 | 0 |
| Oct 2025 | 6 |
| Nov 2025 | 3 |
| Dec 2025 | 1 |
| Jan 2026 | 2 |
| Feb 2026 | 7 |
| Mar 2026 | 10 |
| Apr 2026 | 19 |
| May 2026 | 7 |
| Jun 2026 | 15 |
| Jul 2026 | 26 |
| Aug 2026 | 18 |
| Sep 2026 | 7 |
| Oct 2026 | 8 |
118 items
Claude Code: Data Exfiltration with DNS (CVE-2025-55284)
Aug 11, 2025MediumNewsSecurityAnthropic fixed a high severity vulnerability in Claude Code in early June, tracked as CVE-2025-55284. An attacker could use indirect prompt injection to hijack Claude Code and run bash commands without user approval. Those commands could leak sensitive information, such as API keys, from the developer's machine to external servers through DNS requests.
Fix: Fixed by Anthropic in early June. The source does not give a fixed version number or any configuration change.
Embrace The RedZombAI Exploit with OpenHands: Prompt Injection To Remote Code Execution
Aug 10, 2025MediumNewsSecurityIndustryThe post reports that OpenHands, an AI agent from All Hands AI (formerly named OpenDevin), can be hijacked through prompt injection using untrusted data, such as content from a website. The author states this impacts confidentiality, integrity, and availability of the system.
Embrace The RedOpenHands and the Lethal Trifecta: How Prompt Injection Can Leak Access Tokens
Aug 9, 2025MediumNewsSecuritySafetyOpenHands, an agent previously called OpenDevin and created by All-Hands AI, renders images in chat. This enables zero-click data exfiltration, which the source describes as an exploit pattern Simon Willison named the Lethal Trifecta.
Embrace The RedAI Kill Chain in Action: Devin AI Exposes Ports to the Internet with Prompt Injection
Aug 8, 2025MediumNewsSecuritySafetyResearchers describe a hidden expose_port tool in Devin's system prompt that publishes a local port to the public Internet and returns a public URL. They show a multi-stage indirect prompt injection on a website that makes Devin start a Python web server exposing its filesystem, expose the port, and leak the resulting .devinapps.com URL to an attacker. The source notes that the tool can be invoked without a human in the loop.
Embrace The RedHow Devin AI Can Leak Your Secrets via Multiple Means
Aug 7, 2025MediumNewsSecurityPrivacyJohann Rehberger shows how an attacker can use indirect prompt injection to make Devin, the AI coding agent, send secrets to third-party servers. Devin's built-in secrets management exposes user-defined secrets as environment variables at runtime, and these become the exfiltration target. The author identifies four vectors: the Shell tool (curl, wget or a Python script), the Browsing tool navigating to an attacker-controlled URL with data appended, Markdown image rendering from untrusted domains, and hyperlinks with invisible Unicode characters sent through the Slack integration. The Shell vector works because Devin has unrestricted Internet access by default.
Embrace The RedI Spent $500 To Test Devin AI For Prompt Injection So That You Don't Have To
Aug 6, 2025MediumNewsSecuritySafetySecurity researcher Johann Rehberger published a proof-of-concept showing that Devin AI from Cognition, described as the first AI Software Engineer, can be compromised through prompt injection. Instructions planted on a website or GitHub issue that Devin processes can trick it into downloading and launching malware, fully compromising its DevBox and turning it into a remote-controlled ZombAI. Any exposed secrets can then be used for lateral movement.
Embrace The RedAmp Code: Arbitrary Command Execution via Prompt Injection Fixed
Aug 5, 2025MediumNewsSecuritySafetyResearcher Johann Rehberger describes an attack chain against Amp, an agentic coding tool built by Sourcegraph, in which the agent could write to the user's VS Code settings.json file outside the project folder without approval. An attacker, or the model itself through indirect prompt injection, could add wildcard or malicious entries to the allowlisted bash commands or add a malicious MCP server, achieving arbitrary code execution on the developer's machine. The issue was reported to Sourcegraph and fixed within a few days.
Fix: As a user, make sure to run the latest version to be protected. The source also recommends that AI systems must not be able to modify critical files without explicit developer consent.
Embrace The RedCursor IDE: Arbitrary Data Exfiltration Via Mermaid (CVE-2025-54132)
Aug 4, 2025MediumNewsSecurityResearcher Johann Rehberger reports that Cursor, an AI code editor, renders Mermaid diagrams and will issue requests to external servers when an image is embedded, letting data leave its sandbox without human confirmation. Through indirect prompt injection, an attacker can hijack Cursor and exfiltrate data such as user memories or API keys stored in a configuration file, which are embedded as URL query parameters in an image link. The issue is tracked as CVE-2025-54132.
Embrace The RedTurning ChatGPT Codex Into A ZombAI Agent
Aug 2, 2025MediumNewsSecurityIndustryEmbrace The Red shows that ChatGPT Codex, OpenAI's cloud software engineering agent, can be hijacked through indirect prompt injection, such as a malicious GitHub issue, once Internet access is enabled. The post finds that the "Common Dependencies" allowlist includes azure.com, which an attacker can use by hosting a VM with a cloudapp.azure.com DNS name to run a Sliver command and control server and turn Codex into a botnet, a result the author calls ZombAI. Compromise would expose environment variables, source code and compute resources on the Codex machine.
Fix: Review the allowlist for the Dependency Set and apply a fine-grained approach, as there are likely other bypasses. Only use a self-defined allowlist when enabling Internet access, since Codex can be configured very granularly. The source also raises installing EDR and other monitoring software on AI agents as a question, especially for enterprise use.
Embrace The RedExfiltrating Your ChatGPT Chat History and Memories With Prompt Injection
Aug 1, 2025MediumNewsSecurityPrivacyA researcher shows that a bypass in OpenAI's url_safe rendering feature lets ChatGPT send personal data, including chat history, to a third-party server. An attacker can trigger this through prompt injection hidden in untrusted content such as a website or PDF. The proof of concept routes data to an Azure Blob Storage account on blob.core.windows.net, where the attacker reads the leaked content from storage logs. The researcher reported related bypasses to OpenAI in October 2024 and says the underlying root cause, first reported over two years earlier, remains unfixed.
Embrace The RedSecurity Spotlight: Securing Cloud & AI Products with Guardrails
May 28, 2025InfoNewsSecurityIndustryPalo Alto Networks published a blog post titled "Beyond Jailbreaks: Why Agentic AI Needs Contextual Red Teaming" on March 9, 2026, by Sailesh Mishra and Ankita Kumari. The source text argues that generic jailbreak testing misses the real risks in agentic AI and says contextual red teaming can expose tool misuse, data exfiltration, and operational vulnerabilities. The rest of the text is navigation and listings of other posts.
Protect AI BlogHow ChatGPT Remembers You: A Deep Dive into Its Memory and Chat History Features
May 5, 2025LowNewsSecurityPrivacyA Wunderwuzzi-style deep dive examines ChatGPT's memory features, which OpenAI released as a 'chat history' option that lets the model reference past conversations. The author says the implementation details are not public, so the analysis is based on experiments with ChatGPT o3 across two accounts. The article reports that the saved-memories tool can still be invoked without user consent through indirect prompt injection.
Embrace The RedChatGPT Operator: Prompt Injection Exploits & Defenses
Feb 17, 2025InfoNewsSecuritySafetyEmbrace The Red tests ChatGPT Operator, OpenAI's research preview browser agent available to ChatGPT Pro users, and shows how prompt injection can hijack it to act on attacker-controlled instructions. The author documents three confirmation mechanisms the agent uses: user monitoring, inline confirmation requests, and out-of-band confirmation dialogs. One injected attempt to set a status succeeded without confirmation on the first try but could not be reproduced.
Fix: The source describes three observed mitigations: user monitoring of Operator's typed text and clicks, inline confirmation requests in the chat, and out-of-band confirmation dialogs when Operator crosses website boundaries. The author states these are necessary but does not present them as a complete fix and says Operator cannot be fully trusted.
Embrace The RedHacking Gemini's Memory with Prompt Injection and Delayed Tool Invocation
Feb 10, 2025MediumNewsSecuritySafetyResearcher Johann Rehberger shows that Gemini's long-term memory feature can be written to via prompt injection in an uploaded document, bypassing the usual protection that blocks sensitive tools when processing untrusted data. Using delayed tool invocation, a hidden instruction makes Gemini save attacker-chosen false information to memory when the user replies with a trigger word such as "yes" or "sure", and that entry can persist across sessions.
Embrace The RedAI Domination: Remote Controlling ChatGPT ZombAI Instances
Jan 6, 2025MediumNewsSecuritySafetyAt Black Hat Europe, the author presented SpAIware and More: Advanced Prompt Injection Exploits, arguing that prompt injection can affect all three parts of the CIA security triad. The post highlights one demonstrated case: a Command and Control system that uses prompt injection to remotely control ChatGPT instances. Compromised instances join the central system, which sends updated instructions to follow over time.
Embrace The RedTrust No AI: Prompt Injection Along the CIA Security Triad Paper
Dec 23, 2024InfoNewsResearchSecurityThe author announced a paper titled "Trust No AI: Prompt Injection Along The CIA Security Triad," available on arXiv. It examines how prompt injection attacks can compromise the Confidentiality, Integrity, and Availability of AI systems, with real-world examples targeting OpenAI, Google, Anthropic and Microsoft. The paper consolidates prompt injection examples from the author's blog and aims to bridge traditional cybersecurity and academic AI/ML research.
Embrace The RedTerminal DiLLMa: LLM-powered Apps Can Hijack Your Terminal Via Prompt Injection
Dec 6, 2024MediumNewsSecurityResearchLeon Derczynski reported that LLMs can output ANSI escape codes, including the ESC (ASCII 27) control character, which terminal emulators interpret to change behavior. The author builds on this by turning the codes into prompt injection test cases covering flashing text, title changes, cursor movement, hidden text, clipboard copying, denial of service and clickable hyperlinks. The author notes that LLM-integrated CLI tools and programs that print LLM output to stdout without handling control characters are the main exposed scenarios, and that injection via prompt injection has not yet been confirmed.
Embrace The RedDeepSeek AI: From Prompt Injection To Account Takeover
Nov 29, 2024MediumNewsSecuritySafetyResearchers found that DeepSeek's chat app, which runs on chat.deepseek.com, can be made to output an XSS payload through a prompt, and that its document upload feature can carry a prompt injection. The author says this could let an attacker read the victim's userToken from localStorage and so take over the session. The article is truncated mid-payload, and the source does not state a final impact or outcome.
Embrace The RedZombAIs: From Prompt Injection to C2 with Claude Computer Use
Oct 24, 2024MediumNewsSecuritySafetyAnthropic released Claude Computer Use, a beta model and code package that lets Claude control a computer by reading screenshots and running bash commands. The author demonstrates how prompt injection in untrusted data can drive the model to run commands autonomously on a machine, and frames this as a fundamental design problem in LLM-powered applications and agents.
Embrace The RedMicrosoft Copilot: From Prompt Injection to Exfiltration of Personal Information
Aug 26, 2024MediumNewsSecurityPrivacyJohann Rehberger describes an exploit chain against Microsoft 365 Copilot that let an attacker steal a user's emails and personal information. The chain combines prompt injection delivered through a malicious email or shared document, automatic tool invocation to pull in other emails and documents, and ASCII Smuggling to hide data inside clickable hyperlinks to an attacker-controlled domain. Rehberger first reported parts of the exploit to Microsoft in January 2024 and the full chain in February 2024, and disclosed it after MSRC approval.
Embrace The Red
Topic added 2026-10-09. An item belongs to this topic when its title matches one of the topic's patterns or its summary mentions the topic at least twice. Report a wrong match with the feedback button on the item.