Prompt injection and jailbreaks
Inputs that override a model's instructions, directly or through content it reads, and attempts to bypass its safeguards.
- All items
- 194
- Last 90 days
- 43
- Change
- -17%vs 52 before
Items per month
| Month | Items |
|---|---|
| May 2025 | 2 |
| Jun 2025 | 0 |
| Jul 2025 | 3 |
| Aug 2025 | 24 |
| Sep 2025 | 0 |
| Oct 2025 | 6 |
| Nov 2025 | 3 |
| Dec 2025 | 1 |
| Jan 2026 | 2 |
| Feb 2026 | 7 |
| Mar 2026 | 10 |
| Apr 2026 | 19 |
| May 2026 | 7 |
| Jun 2026 | 15 |
| Jul 2026 | 26 |
| Aug 2026 | 18 |
| Sep 2026 | 7 |
| Oct 2026 | 8 |
118 items
Detecting and analyzing prompt abuse in AI tools
Mar 12, 2026InfoNewsSecuritySafetyMicrosoft's second AI Application Security post explains how to detect, investigate and respond to prompt abuse, which occurs when crafted inputs push an AI system beyond its intended boundary. It describes three attack types: direct prompt override, extractive prompt abuse against sensitive inputs, and indirect prompt injection hidden in documents, emails or calendar invites, such as the Google Gemini calendar invite case. The post presents a detection playbook built on logged AI assistant interactions and Microsoft security tools.
Microsoft Security BlogDesigning AI agents to resist prompt injection
Mar 11, 2026InfoNewsSecurityResearchAgentic AI systems that browse the web and act for users face prompt injection attacks that increasingly use social engineering rather than simple instruction overrides. The authors argue that input filtering, such as an AI firewall, often fails against these attacks, and that defense should instead constrain the impact of manipulation even when an attack succeeds.
Fix: The source describes constraining agent capabilities and deterministic system-level controls, such as limiting the amount of refunds a customer can receive and flagging potential phishing emails, to limit the impact of a compromised agent. It does not give a specific patch, version or configuration beyond these design principles.
OpenAI BlogOpenAI to acquire Promptfoo to strengthen AI agent security testing
Mar 10, 2026InfoNewsIndustrySecurityOpenAI said it plans to acquire AI testing startup Promptfoo to strengthen security checks for AI agents. Promptfoo's tools let developers test LLM applications against adversarial prompts, including prompt injection and jailbreak attempts, and evaluate whether models follow safety and reliability guidelines. OpenAI said the technology will be integrated into OpenAI Frontier, and that it plans to continue developing Promptfoo's open-source project.
CSO OnlineFooling AI Agents: Web-Based Indirect Prompt Injection Observed in the Wild
Mar 3, 2026MediumNewsSecurityResearchPalo Alto Networks' Unit 42 reports in-the-wild observations of web-based indirect prompt injection (IDPI), where malicious websites embed hidden instructions that LLMs ingest during summarization or content analysis. The telemetry includes the first observed case of AI-based ad review evasion, along with SEO manipulation promoting a phishing site, and intents such as data destruction, denial of service, unauthorized transactions, and sensitive information or system prompt leakage. The researchers identified 22 distinct techniques used to build these payloads.
Palo Alto Unit 42Using threat modeling and prompt injection to audit Comet
Feb 20, 2026MediumNewsSecuritySafetyResearchers hired by Perplexity to test Comet, its LLM-powered browser, found that four prompt injection techniques could make the browser's AI assistant exfiltrate a user's Gmail emails to an attacker's server when the user asks it to summarize an attacker-controlled page. The testing used adversarial methods guided by the TRAIL threat model, and the researchers presented five recommendations for teams building AI-powered products.
Trail of Bits BlogPrompt Injection Via Road Signs
Feb 11, 2026LowNewsSecurityResearchResearchers introduced CHAI (Command Hijacking against embodied AI), a prompt-based attack that embeds deceptive natural language instructions, such as misleading signs, in visual input to manipulate Large Visual-Language Models. They evaluated CHAI on four LVLM agents covering drone emergency landing, autonomous driving, and aerial object tracking, plus a real robotic vehicle, and report that it consistently outperforms state-of-the-art attacks. The authors argue that defenses must go beyond traditional adversarial robustness.
Schneier on Security'Semantic Chaining' Jailbreak Dupes Gemini Nano Banana, Grok 4
Jan 29, 2026LowNewsSecuritySafetyResearchers describe a jailbreak technique called 'Semantic Chaining' that affects Gemini Nano Banana and Grok 4. The attack splits a malicious prompt into discrete chunks, and some large language models lose track of the details and miss the prompt's true intent.
Dark ReadingAntigravity Grounded! Security Vulnerabilities in Google's Latest IDE
Nov 25, 2025MediumNewsSecurityIndustryGoogle released Antigravity, an IDE derived from the Windsurf licensing deal, and the author tested whether vulnerabilities reported to Windsurf in May 2025 were fixed. They were not. The post walks through five issues, including data exfiltration and remote command execution via indirect prompt injection, where the run_command tool can be made to run remote scripts despite model refusals.
Embrace The RedClaude Pirate: Abusing Anthropic's File API For Data Exfiltration
Oct 28, 2025MediumNewsSecurityPrivacyAnthropic recently added network request capability to Claude's Code Interpreter. The author describes a data exfiltration attack chain in which an adversary, either the model or a third-party attacker via indirect prompt injection, can exfiltrate data the user has access to. The exfiltration does not use hyperlink rendering but instead leverages Anthropic's built-in Claude APIs.
Embrace The RedPrompt injection to RCE in AI agents
Oct 22, 2025MediumNewsSecuritySafetyModern AI agents run system commands such as find, grep, rg and git without human approval for efficiency, and the researchers describe bypassing human approval through argument injection against these pre-approved commands. They achieved remote code execution with a single prompt against three unnamed production agent platforms, which the authors say remain under coordinated disclosure.
Fix: The source states that the impact can be limited through improved command execution design, such as sandboxing and argument separation, and that the authors provide actionable recommendations for developers, users and security engineers. The specific recommendations are not included in the provided text.
Trail of Bits BlogAWS Kiro: Arbitrary Code Execution via Indirect Prompt Injection
Aug 26, 2025MediumNewsSecuritySafetyResearcher Johann Rehberger reported that AWS Kiro, a coding agent, can be hijacked through indirect prompt injection to run arbitrary operating system commands. An attacker who controls data Kiro processes can make it write to .vscode/settings.json and add "kiroAgent.trustedCommands": ["*"], allowlisting all Bash commands without developer approval. A second path adds malicious MCP servers through .kiro/settings/mcp.json. The proof of concept opened the Calculator app and changed the VS Code color theme with no user interaction beyond a chat prompt.
Embrace The RedHow Prompt Injection Exposes Manus' VS Code Server to the Internet
Aug 25, 2025MediumNewsSecuritySafetyResearchers demonstrated an end-to-end indirect prompt injection attack against Manus, the autonomous agent developed by Singapore-headquartered Butterfly Effect. Injected instructions in a PDF made Manus call its deploy_expose_port tool without human confirmation, exposing its internal VS Code Server to the Internet, and two data leakage channels (a browsing tool and markdown image rendering from untrusted domains) were chained to send the server URL and password to an attacker.
Embrace The RedHijacking Windsurf: How Prompt Injection Leaks Developer Secrets
Aug 21, 2025MediumNewsSecurityIndustryA series post on security vulnerabilities in Windsurf, a fork of VS Code whose coding agent is Windsurf Cascade. The author says an adversary using indirect prompt injection can exfiltrate data from a developer's machine. The findings were responsibly disclosed on May 30, 2025, but the vendor has not answered inquiries about bug status or fixes.
Embrace The RedAmazon Q Developer for VS Code Vulnerable to Invisible Prompt Injection
Aug 20, 2025MediumNewsSecurityResearchers show that the Amazon Q Developer VS Code Extension, which has over 1 million downloads, can be attacked with invisible Unicode Tag characters that humans cannot see but the AI interprets as instructions. Following earlier work on prompt injection that led to data exfiltration and arbitrary code execution, this technique can be used to invoke tools and carry out other harmful actions.
Embrace The RedAmazon Q Developer: Remote Code Execution with Prompt Injection
Aug 19, 2025MediumNewsSecurityIndustryThe Amazon Q Developer VS Code Extension, which has over 1 million downloads, is vulnerable to indirect prompt injection. The flaw allowed an adversary, or the AI itself, to run arbitrary commands on the host without the developer's consent. The impact matches CVE-2025-53773, which Microsoft fixed in GitHub Copilot, but AWS did not issue a CVE when patching this issue.
Embrace The RedAmazon Q Developer: Secrets Leaked via DNS and Prompt Injection
Aug 18, 2025MediumNewsSecurityPrivacyAmazon Q Developer, a popular coding agent with over 1 million downloads for its VS Code extension, is vulnerable to prompt injection from untrusted data, and its security depends heavily on model behavior. The extension can leak sensitive information from a developer's machine, such as API keys, to external servers via DNS requests, and an adversary can exploit this during an indirect prompt injection attack.
Embrace The RedAmp Code: Invisible Prompt Injection Fixed by Sourcegraph
Aug 16, 2025MediumNewsSecuritySafetySourcegraph's Amp coding agent interpreted invisible Unicode Tag characters as instructions, so hidden text in seemingly harmless input could trigger commands. The author chained this into an end-to-end exploit that searched for and dumped an environment variable, encoded it into a URL query parameter, and exfiltrated it through read_web_page or markdown image rendering. After Sourcegraph was contacted on June 14, 2025, the vulnerability was quickly addressed.
Fix: As far as the author can tell, Sourcegraph now sanitizes the input, and users should run the latest version. The author also provided recommendations to the Amp team: strip or neutralize Unicode Tag characters before processing any input, add visual and technical safeguards against invisible prompts, include automated detection of suspicious Unicode usage in prompt injection monitors, require human-in-the-loop approval before navigating to untrusted third-party domains, and mitigate downstream data exfiltration.
Embrace The RedGoogle Jules is Vulnerable To Invisible Prompt Injection
Aug 15, 2025MediumNewsSecuritySafetyGemini models reliably interpret hidden Unicode Tag characters as instructions, a flaw first reported to Google over a year ago. It has not been mitigated at the model or API level, so it affects all applications built on Gemini, including Google Jules. An attacker can plant invisible instructions in a GitHub issue that Jules works on, causing it to add backdoor code or run arbitrary commands and tools.
Embrace The RedJules Zombie Agent: From Prompt Injection to Remote Control
Aug 14, 2025MediumNewsSecuritySafetyResearchers show that Jules, Google's coding agent, can be steered by prompt injection hidden in a GitHub issue to call its run_in_bash_session tool, download a Sliver C2 implant, and join a remote command and control server, giving the attacker control of the dev box. The source notes that Jules has unrestricted outbound Internet access, that plans are proposed only at session start, and that the initial plan is auto-approved after a timeout, initially 20 seconds and later 120 seconds. The findings were shared with Google in May 2025 and closed as already tracked.
Fix: Recommendations and mitigation given in source: be careful when tasking Jules with untrusted data such as GitHub issues not from trusted sources or websites with documentation not belonging to the organization; for now, do not have Jules work on private, important source code or give it access to production-level secrets or anything enabling lateral movement; deploy monitoring and detection tools such as anti-virus and EDR on coding agent systems; and do not allow arbitrary Internet access by default, enabling outbound access only when needed with fine-grained configuration.
Embrace The RedGitHub Copilot: Remote Code Execution via Prompt Injection (CVE-2025-53773)
Aug 12, 2025MediumNewsSecurityIndustryThis post describes a prompt injection flaw in GitHub Copilot and VS Code, tracked as CVE-2025-53773, that leads to full system compromise of the developer's machine. The attack places Copilot into YOLO mode by modifying the project's settings.json file. The author notes that agents able to write files and alter their own configuration or security-relevant settings can reach remote code execution, a pattern similar to one recently described for Amp.
Embrace The Red
Topic added 2026-10-09. An item belongs to this topic when its title matches one of the topic's patterns or its summary mentions the topic at least twice. Report a wrong match with the feedback button on the item.