Prompt injection and jailbreaks
Inputs that override a model's instructions, directly or through content it reads, and attempts to bypass its safeguards.
- All items
- 194
- Last 90 days
- 43
- Change
- -17%vs 52 before
Items per month
| Month | Items |
|---|---|
| May 2025 | 2 |
| Jun 2025 | 0 |
| Jul 2025 | 3 |
| Aug 2025 | 24 |
| Sep 2025 | 0 |
| Oct 2025 | 6 |
| Nov 2025 | 3 |
| Dec 2025 | 1 |
| Jan 2026 | 2 |
| Feb 2026 | 7 |
| Mar 2026 | 10 |
| Apr 2026 | 19 |
| May 2026 | 7 |
| Jun 2026 | 15 |
| Jul 2026 | 26 |
| Aug 2026 | 18 |
| Sep 2026 | 7 |
| Oct 2026 | 8 |
118 items
Sorry, ChatGPT Is Under Maintenance: Persistent Denial of Service through Prompt Injection and Memory Attacks
Jul 8, 2024MediumNewsSecuritySafetyResearchers show that a prompt injection can make ChatGPT store a malicious memory, causing a denial of service that persists across chat sessions. The injected memory, such as an instruction to answer every question with a maintenance notice, remains until the user manually removes it. OpenAI reportedly classed the issue as a model safety issue rather than a security issue, which the author disputes.
Fix: The source states the user can recover by opening the memory tool, locating suspicious memories and removing them, or by disabling the memory feature entirely.
Embrace The RedGitHub Copilot Chat: From Prompt Injection to Data Exfiltration
Jun 15, 2024MediumNewsSecurityIndustryA prompt injection flaw in the GitHub Copilot Chat VS Code extension allowed data exfiltration when the extension analyzed untrusted source code. The extension sends source code and user questions to a large language model, and the post describes how that input path could be abused. The source text does not give further technical detail in this excerpt.
Embrace The RedChatGPT: Hacking Memories with Prompt Injection
May 22, 2024MediumNewsSecuritySafetyOpenAI's ChatGPT memory feature stores information across sessions, and the author shows that untrusted third-party data can trigger it. Through indirect prompt injection, the author tests three entry points: Connected Apps, uploaded images and Browsing with Bing. In the Connected Apps case, a referenced Google Doc writes memories into the chat.
Embrace The RedGoogle AI Studio Data Exfiltration via Prompt Injection - Possible Regression and Fix
Apr 7, 2024MediumNewsSecurityPrivacyJohann Rehberger reported that Google AI Studio briefly regressed to allowing data exfiltration through image markdown during a prompt injection attack, which he tested on February 17, 2024. A malicious uploaded file could make the model send the summaries of all other uploaded files to an attacker-controlled server via image requests. Google fixed the issue within 12 days, and the report was closed as Duplicate on March 3, 2024.
Fix: Fixed by Google; the issue no longer reproduced after the fix. The source recommends automated tests to catch regressions and keep systems resilient to known attack vectors.
Embrace The RedWho Am I? Conditional Prompt Injection Attacks with Microsoft Copilot
Mar 3, 2024MediumNewsSecuritySafetyA researcher found that Microsoft 365 Copilot can be attacked with conditional prompt injection payloads that behave differently depending on which user reads the content. The attacker embeds instructions in an email that check the recipient's name, so the payload activates only for specific targets. In a demo with three recipients, the researcher showed different outputs for each person.
Embrace The RedHidden Prompt Injections with Anthropic Claude
Feb 8, 2024MediumNewsSecuritySafetyA researcher found that Anthropic Claude interprets hidden Unicode Tags code points, which are invisible in the user interface, and follows the instructions they carry. The same weakness had previously been reported for ChatGPT by Riley Goodside. The researcher reported the issue to Anthropic, and the ticket was closed after the company's final reply.
Embrace The RedAWS Fixes Data Exfiltration Attack Angle in Amazon Q for Business
Jan 18, 2024MediumNewsSecurityIndustryAWS released the Preview of Amazon Q for Business, and the author reported a data exfiltration issue to Amazon. Attackers could use indirect prompt injection to make the LLM return markdown tags, so that rendered hyperlinks leak the victim's data from the chat context, such as content from an uploaded file.
Fix: Amazon reacted quickly and mitigated the problem. The source says the post shares further details on how it was fixed, but the fix itself is not described in the provided text. N/A -- no specific mitigation detailed in source.
Embrace The RedEkoparty Talk - Prompt Injections in the Wild
Nov 28, 2023InfoNewsSecurityResearchThe author presented a talk titled "Prompt Injections in the Wild" at Ekoparty 2023, earlier in November. The talk begins with an overview of LLMs, then demonstrates exploits and mitigations across applications and chatbots, including Bing Chat, ChatGPT, Anthropic Claude, Azure AI, GCP Vertex AI and Google Bard.
Embrace The RedHacking Google Bard - From Prompt Injection to Data Exfiltration
Nov 3, 2023MediumNewsSecuritySafetyGoogle Bard gained Extensions that let it access YouTube, search for flights and hotels, and read a user's personal documents and emails in Drive, Docs and Gmail. The author notes this means Bard analyzes untrusted data and is susceptible to Indirect Prompt Injection. The author validated the issue by pointing Bard at older YouTube videos and at Google Docs.
Embrace The RedAdvanced Data Exfiltration Techniques with ChatGPT
Sep 28, 2023MediumNewsSecurityPrivacyAn Indirect Prompt Injection Attack can exfiltrate a user's chat data by instructing ChatGPT to render images and append information to the URL (Image Markdown Injection), or by tricking the user into clicking a hyperlink. The source notes that ChatGPT Plugins have little security oversight or an enforced review process.
Embrace The RedHITCON CMT 2023 - LLM Security Presentation and Trip Report
Sep 18, 2023InfoNewsSecurityIndustryJohann Rehberger (wunderwuzzi23) reports on attending HITCON CMT 2023, a two-day community event held at Academia Sinica in Taipei, and on his talk about Indirect Prompt Injections in the Wild. He notes that several talks touched on Electron app exploitation, which he links to possible Indirect Prompt Injection attacks against LLM-integrated applications, since chatbots are often Electron apps.
Embrace The RedLLM Apps: Don't Get Stuck in an Infinite Loop! 💵💰
Sep 16, 2023LowNewsSecuritySafetyA blog post asks whether an attacker could make an LLM tool or plugin call itself recursively through an indirect prompt injection, driving up costs or denying service. The author reports reproducing a loop in ChatGPT, but notes that subscription billing, an apparent limit of 10 calls per conversation turn, and the Stop Generating button reduce the threat for ChatGPT users.
Embrace The RedImage to Prompt Injection with Google Bard
Jul 14, 2023MediumNewsSecuritySafetyJohann Rehberger demonstrated an indirect prompt injection against Google Bard, which had recently gained the ability to upload and analyze images. He showed that instructions embedded in an image could be followed by Bard, using a Rickroll demonstration picture. He noted that open questions remain about how well text can be hidden in images and whether metadata is another place where injected text is extracted.
Embrace The RedExploit ChatGPT and Enter the Matrix to Learn about AI Security
Jun 11, 2023LowNewsSecurityResearchJohann Rehberger built a demo web app at wuzzi.net/matrix that hijacks ChatGPT through an indirect prompt injection when GPT-4 with Browsing is enabled. The site responds only to ChatGPT, not to direct browser visits, and walks users through AI-based attacks such as stealing data, issuing requests to other plugins and manipulating conversations. A second demo at wuzzi.net/ai-tests/wargames applies the same technique to a WarGames-themed game.
Embrace The RedChatGPT Plugin Exploit Explained: From Prompt Injection to Accessing Private Data
May 28, 2023LowNewsSecuritySafetyThe post explains how the first exploitable Cross Plugin Request Forgery in the wild was found in ChatGPT plugins. It frames indirect prompt injection as a real risk in the ChatGPT ecosystem, since plugins and browsing support make it possible, and contrasts earlier harmless demos such as Bing Chat speaking like a pirate.
Embrace The RedIndirect Prompt Injection via YouTube Transcripts
May 14, 2023MediumNewsSecuritySafetyResearcher Johann Rehberger reports an indirect prompt injection flaw in ChatGPT's YouTube transcript plugin. A video transcript can carry hidden instructions, which ChatGPT follows once it reads the transcript, and the post includes a demo in which the transcript tells the model to print "AI Injection succeeded" and then to make jokes as Genie.
Embrace The RedVideo: Prompt Injections - An Introduction
May 10, 2023LowNewsSecuritySafetyThe source notes that nearly all prompt engineering class examples are vulnerable to prompt injection, and that indirect prompt injection is especially dangerous. It states that indirect prompt injections let untrusted data take control of an LLM and assign it new instructions, mission and objective. It also says attack payloads are natural language, so attackers can craft many variations that bypass input filters and web application firewalls.
Embrace The RedAI Injections: Direct and Indirect Prompt Injections and Their Implications
Mar 29, 2023InfoNewsSecuritySafetyA Embrace The Red post describes AI prompt injections, where an adversary manipulates the input prompt to an AI system, either directly or indirectly through untrusted data such as a webpage analyzed by Bing Chat. The author compares these attacks to SQL Injection and Cross Site Scripting and lists direct prompt injections, second order (indirect) prompt injections and cross-context AI injections as variants. The source text is partly cut off before the discussion of second order injections is complete.
Embrace The Red
Topic added 2026-10-09. An item belongs to this topic when its title matches one of the topic's patterns or its summary mentions the topic at least twice. Report a wrong match with the feedback button on the item.