Prompt injection and jailbreaks
Inputs that override a model's instructions, directly or through content it reads, and attempts to bypass its safeguards.
- All items
- 194
- Last 90 days
- 43
- Change
- -17%vs 52 before
Items per month
| Month | Items |
|---|---|
| May 2025 | 2 |
| Jun 2025 | 0 |
| Jul 2025 | 3 |
| Aug 2025 | 24 |
| Sep 2025 | 0 |
| Oct 2025 | 6 |
| Nov 2025 | 3 |
| Dec 2025 | 1 |
| Jan 2026 | 2 |
| Feb 2026 | 7 |
| Mar 2026 | 10 |
| Apr 2026 | 19 |
| May 2026 | 7 |
| Jun 2026 | 15 |
| Jul 2026 | 26 |
| Aug 2026 | 18 |
| Sep 2026 | 7 |
| Oct 2026 | 8 |
194 items
Hacking Gemini's Memory with Prompt Injection and Delayed Tool Invocation
Feb 10, 2025MediumNewsSecuritySafetyResearcher Johann Rehberger shows that Gemini's long-term memory feature can be written to via prompt injection in an uploaded document, bypassing the usual protection that blocks sensitive tools when processing untrusted data. Using delayed tool invocation, a hidden instruction makes Gemini save attacker-chosen false information to memory when the user replies with a trigger word such as "yes" or "sure", and that entry can persist across sessions.
Embrace The RedAI Domination: Remote Controlling ChatGPT ZombAI Instances
Jan 6, 2025MediumNewsSecuritySafetyAt Black Hat Europe, the author presented SpAIware and More: Advanced Prompt Injection Exploits, arguing that prompt injection can affect all three parts of the CIA security triad. The post highlights one demonstrated case: a Command and Control system that uses prompt injection to remotely control ChatGPT instances. Compromised instances join the central system, which sends updated instructions to follow over time.
Embrace The RedTrust No AI: Prompt Injection Along the CIA Security Triad Paper
Dec 23, 2024InfoNewsResearchSecurityThe author announced a paper titled "Trust No AI: Prompt Injection Along The CIA Security Triad," available on arXiv. It examines how prompt injection attacks can compromise the Confidentiality, Integrity, and Availability of AI systems, with real-world examples targeting OpenAI, Google, Anthropic and Microsoft. The paper consolidates prompt injection examples from the author's blog and aims to bridge traditional cybersecurity and academic AI/ML research.
Embrace The RedTerminal DiLLMa: LLM-powered Apps Can Hijack Your Terminal Via Prompt Injection
Dec 6, 2024MediumNewsSecurityResearchLeon Derczynski reported that LLMs can output ANSI escape codes, including the ESC (ASCII 27) control character, which terminal emulators interpret to change behavior. The author builds on this by turning the codes into prompt injection test cases covering flashing text, title changes, cursor movement, hidden text, clipboard copying, denial of service and clickable hyperlinks. The author notes that LLM-integrated CLI tools and programs that print LLM output to stdout without handling control characters are the main exposed scenarios, and that injection via prompt injection has not yet been confirmed.
Embrace The RedDeepSeek AI: From Prompt Injection To Account Takeover
Nov 29, 2024MediumNewsSecuritySafetyResearchers found that DeepSeek's chat app, which runs on chat.deepseek.com, can be made to output an XSS payload through a prompt, and that its document upload feature can carry a prompt injection. The author says this could let an attacker read the victim's userToken from localStorage and so take over the session. The article is truncated mid-payload, and the source does not state a final impact or outcome.
Embrace The RedZombAIs: From Prompt Injection to C2 with Claude Computer Use
Oct 24, 2024MediumNewsSecuritySafetyAnthropic released Claude Computer Use, a beta model and code package that lets Claude control a computer by reading screenshots and running bash commands. The author demonstrates how prompt injection in untrusted data can drive the model to run commands autonomously on a machine, and frames this as a fundamental design problem in LLM-powered applications and agents.
Embrace The RedCVE-2024-48142: Monica ChatGPT AI Assistant prompt injection in chatbox exposes chat data
Oct 24, 2024HighVulnerabilitySecurityCVE-2024-48142CVE-2024-48142 describes a prompt injection vulnerability in the chatbox of Butterfly Effect Limited's Monica ChatGPT AI Assistant v2.4.0. A crafted message allows attackers to access and exfiltrate all previous and subsequent chat data between the user and the AI assistant. The NVD had not yet provided an assessment, and the weakness is mapped to CWE-77.
NVD/CVE DatabaseCVE-2024-48140: Monica Your AI Copilot prompt injection in chatbox exposes chat data
Oct 24, 2024HighVulnerabilitySecurityCVE-2024-48140CVE-2024-48140 describes a prompt injection vulnerability in the chatbox of Butterfly Effect Limited's Monica Your AI Copilot powered by ChatGPT4, version 6.3.0. A crafted message lets an attacker access and exfiltrate all previous and subsequent chat data between the user and the AI assistant. The source lists CWE-77 (Command Injection) as the weakness, and NVD has not yet provided an assessment.
NVD/CVE DatabaseCVE-2024-48145: Netangular Technologies ChatNet AI prompt injection in chatbox exposes chat data
Oct 24, 2024CriticalVulnerabilitySecurityCVE-2024-48145CVE-2024-48145 describes a prompt injection vulnerability in the chatbox of Netangular Technologies ChatNet AI, Version v1.0. A crafted message allows attackers to access and exfiltrate all previous and subsequent chat data between the user and the AI assistant. The NVD assessment is not yet provided, and the weakness is classified as CWE-77.
NVD/CVE DatabaseCVE-2024-48144: Fusion Chat AI Assistant prompt injection exposes chat data
Oct 24, 2024CriticalVulnerabilitySecurityCVE-2024-48144CVE-2024-48144 describes a prompt injection vulnerability in the chatbox of Fusion Chat Chat AI Assistant Ask Me Anything v1.2.4.0. A crafted message lets attackers access and exfiltrate all previous and subsequent chat data between the user and the AI assistant. NVD has not yet provided an assessment, and the record is linked to CWE-77 (Command Injection).
NVD/CVE DatabaseCVE-2024-48141: Zhipu AI CodeGeeX prompt injection in chatbox exposes chat data
Oct 24, 2024HighVulnerabilitySecurityCVE-2024-48141CVE-2024-48141 is a prompt injection vulnerability in the chatbox of Zhipu AI CodeGeeX v2.17.0. A crafted message lets an attacker access and exfiltrate all previous and subsequent chat data between the user and the AI assistant. The weakness is catalogued as CWE-77, Improper Neutralization of Special Elements used in a Command ('Command Injection'), and NVD published the entry on 10/24/2024.
NVD/CVE DatabaseCVE-2024-48139: Blackbox AI prompt injection in chatbox exposes chat data
Oct 24, 2024HighVulnerabilitySecurityCVE-2024-48139CVE-2024-48139 is a prompt injection vulnerability in the chatbox of Blackbox AI v1.3.95. A crafted message lets an attacker access and exfiltrate all previous and subsequent chat data between the user and the AI assistant. The NVD assessment has not yet been provided, and the source maps the weakness to CWE-77 (Command Injection).
NVD/CVE DatabaseMicrosoft Copilot: From Prompt Injection to Exfiltration of Personal Information
Aug 26, 2024MediumNewsSecurityPrivacyJohann Rehberger describes an exploit chain against Microsoft 365 Copilot that let an attacker steal a user's emails and personal information. The chain combines prompt injection delivered through a malicious email or shared document, automatic tool invocation to pull in other emails and documents, and ASCII Smuggling to hide data inside clickable hyperlinks to an attacker-controlled domain. Rehberger first reported parts of the exploit to Microsoft in January 2024 and the full chain in February 2024, and disclosed it after MSRC approval.
Embrace The RedSorry, ChatGPT Is Under Maintenance: Persistent Denial of Service through Prompt Injection and Memory Attacks
Jul 8, 2024MediumNewsSecuritySafetyResearchers show that a prompt injection can make ChatGPT store a malicious memory, causing a denial of service that persists across chat sessions. The injected memory, such as an instruction to answer every question with a maintenance notice, remains until the user manually removes it. OpenAI reportedly classed the issue as a model safety issue rather than a security issue, which the author disputes.
Fix: The source states the user can recover by opening the memory tool, locating suspicious memories and removing them, or by disabling the memory feature entirely.
Embrace The RedGitHub Copilot Chat: From Prompt Injection to Data Exfiltration
Jun 15, 2024MediumNewsSecurityIndustryA prompt injection flaw in the GitHub Copilot Chat VS Code extension allowed data exfiltration when the extension analyzed untrusted source code. The extension sends source code and user questions to a large language model, and the post describes how that input path could be abused. The source text does not give further technical detail in this excerpt.
Embrace The RedCVE-2024-5184: EmailGPT prompt injection that leaks system prompts and runs unwanted prompts
Jun 5, 2024MediumVulnerabilitySecuritySafetyCVE-2024-5184CVE-2024-5184 affects the EmailGPT service, which relies on an API service that lets a malicious user inject a direct prompt and take over the service logic. An attacker can force the AI service to leak its hard-coded system prompts and/or execute unwanted prompts, and any individual with access to the service can exploit it.
NVD/CVE DatabaseChatGPT: Hacking Memories with Prompt Injection
May 22, 2024MediumNewsSecuritySafetyOpenAI's ChatGPT memory feature stores information across sessions, and the author shows that untrusted third-party data can trigger it. Through indirect prompt injection, the author tests three entry points: Connected Apps, uploaded images and Browsing with Bing. In the Connected Apps case, a referenced Google Doc writes memories into the chat.
Embrace The RedGoogle AI Studio Data Exfiltration via Prompt Injection - Possible Regression and Fix
Apr 7, 2024MediumNewsSecurityPrivacyJohann Rehberger reported that Google AI Studio briefly regressed to allowing data exfiltration through image markdown during a prompt injection attack, which he tested on February 17, 2024. A malicious uploaded file could make the model send the summaries of all other uploaded files to an attacker-controlled server via image requests. Google fixed the issue within 12 days, and the report was closed as Duplicate on March 3, 2024.
Fix: Fixed by Google; the issue no longer reproduced after the fix. The source recommends automated tests to catch regressions and keep systems resilient to known attack vectors.
Embrace The RedWho Am I? Conditional Prompt Injection Attacks with Microsoft Copilot
Mar 3, 2024MediumNewsSecuritySafetyA researcher found that Microsoft 365 Copilot can be attacked with conditional prompt injection payloads that behave differently depending on which user reads the content. The attacker embeds instructions in an email that check the recipient's name, so the payload activates only for specific targets. In a demo with three recipients, the researcher showed different outputs for each person.
Embrace The RedHidden Prompt Injections with Anthropic Claude
Feb 8, 2024MediumNewsSecuritySafetyA researcher found that Anthropic Claude interprets hidden Unicode Tags code points, which are invisible in the user interface, and follows the instructions they carry. The same weakness had previously been reported for ChatGPT by Riley Goodside. The researcher reported the issue to Anthropic, and the ticket was closed after the company's final reply.
Embrace The Red
Topic added 2026-10-09. An item belongs to this topic when its title matches one of the topic's patterns or its summary mentions the topic at least twice. Report a wrong match with the feedback button on the item.