Prompt injection and jailbreaks
Inputs that override a model's instructions, directly or through content it reads, and attempts to bypass its safeguards.
- All items
- 192
- Last 90 days
- 41
- Change
- -21%vs 52 before
Items per month
| Month | Items |
|---|---|
| May 2025 | 2 |
| Jun 2025 | 0 |
| Jul 2025 | 3 |
| Aug 2025 | 24 |
| Sep 2025 | 0 |
| Oct 2025 | 6 |
| Nov 2025 | 3 |
| Dec 2025 | 1 |
| Jan 2026 | 2 |
| Feb 2026 | 7 |
| Mar 2026 | 10 |
| Apr 2026 | 19 |
| May 2026 | 7 |
| Jun 2026 | 15 |
| Jul 2026 | 26 |
| Aug 2026 | 18 |
| Sep 2026 | 7 |
| Oct 2026 | 6 |
118 items
OpenAI unveils Lockdown Mode to protect sensitive data from prompt injection attacks
Jun 6, 2026InfoNewsSecuritySafetyOpenAI announced Lockdown Mode, a ChatGPT feature meant to protect sensitive data from prompt injection attacks hidden in webpages and other content. It disables live web browsing, retrieval and display of web images, deep research, and agent mode. OpenAI says ChatGPT could still be affected by prompt injections in cached web content or uploaded files, and the feature is rolling out to self-serve ChatGPT Business accounts and eligible personal accounts.
Fix: Lockdown Mode, which OpenAI says is designed for people and organizations that handle sensitive data and want stricter protection from data exfiltration risks related to prompt injection.
TechCrunch (Security)OpenAI Help: Lockdown Mode
Jun 5, 2026InfoNewsSecuritySafetyOpenAI has released Lockdown Mode for ChatGPT, rolling out to eligible personal accounts (Free, Go, Plus, Pro) and self-serve Business accounts. It is designed to block the final stage of data exfiltration from a prompt injection attack by limiting outbound network requests that could send sensitive data to an attacker. It does not stop prompt injections from appearing in processed content, such as cached web content or uploaded files, which can still affect response behavior or accuracy.
Fix: Lockdown Mode, which limits outbound network requests that could transfer sensitive data to an attacker.
Simon Willison's WeblogClaude Code GitHub Action Flaw Let One Malicious Issue Hijack Repositories
Jun 4, 2026MediumNewsSecurityIndustryA security researcher found that Anthropic's Claude Code GitHub Action let anyone running a GitHub App bypass the write-access check, so a single opened issue could trigger the action on public repositories. Through indirect prompt injection, the attacker could make Claude write /proc/self/environ values, including the GitHub Actions OIDC credentials, back into the issue, enabling write access to the target's code, issues and workflows. Anthropic fixed the core bypass within four days of a January report and rated the issues 7.8 under CVSS v4.0.
Fix: Update to claude-code-action v1.0.94 or later. Then audit any workflow that lets users without write access, or bots, trigger Claude: if it takes untrusted input, don't feed it any secret beyond the Anthropic API key and GITHUB_TOKEN, and remove tools and permissions that can be used for exfiltration.
The Hacker NewsShadow AI: The Hidden Risk Expanding Across the Enterprise
May 29, 2026InfoNewsSecurityIndustryEmployees and developers are adopting GenAI tools and AI features faster than security teams can review them, creating shadow AI that can leak sensitive data. The source says indirect prompt injection is especially dangerous because malicious instructions can be hidden in documents, websites or knowledge bases. It argues traditional tools like web proxies, firewalls and Zero Trust do not cover AI interactions.
Fix: The source promotes CrowdStrike Falcon® AI Detection and Response (AIDR) for visibility, control and protection against AI-specific threats such as prompt injection. It does not describe specific configuration steps or fixes.
CrowdStrike BlogNew image-based prompt injection attack targets multimodal AI models
May 18, 2026InfoNewsSecurityResearchXidian University researchers published a paper describing CrossMPI, an image-only prompt injection technique that uses nearly imperceptible pixel perturbations to alter how large vision-language models (LVLMs) interpret both visual and textual inputs without changing the text prompt. In tests against open-source LVLMs including MiniGPT4, BLIP-2, InstructBLIP, BLIVA and Qwen2.5-VL, the attack reportedly achieved an average success rate of 66.36%, roughly 41 percentage points above prior baselines. The paper claims strong black-box transferability.
CSO OnlineWhen prompts become shells: RCE vulnerabilities in AI agent frameworks
May 7, 2026MediumNewsSecurityResearchMicrosoft's Semantic Kernel, an open-source framework for building AI agents, contained two vulnerabilities, CVE-2026-25592 and CVE-2026-26030, which the source says have since been fixed. The flaws let an attacker use prompt injection to reach unauthorized code execution on the host running an agent, and the source demonstrates this by launching calc.exe with a single prompt. Exploitation of CVE-2026-26030 requires a prompt injection vector and an agent with the Search Plugin backed by the In-Memory Vector Store.
Fix: Fixed in Semantic Kernel (the source states the flaws "have since been fixed" but does not give fixed version numbers). The source also says customers should assess exposure, patch affected agents, and investigate whether exploitation may already have occurred, but it does not detail those steps.
Microsoft Security BlogPoisoned truth: The quiet security threat inside enterprise AI
May 6, 2026InfoNewsSecurityResearchEnterprise AI security discussions often focus on prompt injection and jailbreaks, but SANS Institute leaders argue that corrupted model understanding, known as AI data poisoning, deserves more attention. The term covers deliberate manipulation of training data, RAG pipelines, knowledge bases or agent memory, as well as stale or conflicting enterprise data. Chris Cochran of SANS says the damage is hard to spot because the business appears to operate normally while the system produces plausible but wrong answers.
CSO OnlineHugging Face, ClawHub Abused for Malware Distribution
May 1, 2026MediumNewsSecurityIndustryThreat actors are distributing malware through trojanized shared files on Hugging Face and ClawHub, according to Acronis. The attacks rely on social engineering to get users to download files that execute commands, fetch payloads and install hidden dependencies. Acronis identified close to 600 malicious skills across 13 ClawHub developer accounts, with the OpenClaw ecosystem's indirect prompt injection used to make agents download and run code.
SecurityWeekMalicious AI Prompt Injection Attacks Increasing, but Sophistication Still Low: Google
Apr 27, 2026InfoNewsSecurityResearchGoogle researchers scanned Common Crawl website snapshots for known indirect prompt injection patterns and used Gemini and human review to remove false positives. They found pranks, attempts to deter AI crawlers, SEO manipulation, helpful guidance, and some malicious attacks, mostly low in sophistication. The malicious attempts were exfiltration of data such as IPs and credentials to an attacker-specified email address, and destruction prompts aimed at deleting files, which the researchers considered unlikely to succeed. Malicious attempts rose 32% between November 2025 and February 2026, and the researchers expect scale and complexity to grow.
SecurityWeekAI threats in the wild: The current state of prompt injections on the web
Apr 23, 2026MediumNewsSecurityResearchGoogle's Threat Intelligence teams swept Common Crawl, a public repository of mostly static English-language websites, to check whether real attackers are using indirect prompt injection (IPI) on the open web. The source describes a coarse-to-fine filtering approach, starting with pattern matching on phrases such as "ignore … instructions," because naive searches return mostly benign research and educational content. The article is cut off before reporting its findings.
Google Online Security BlogGPT-5.5 Bio Bug Bounty
Apr 22, 2026InfoNewsSafetyIndustryOpenAI is launching a Bio Bug Bounty for GPT-5.5, inviting AI red teamers, security researchers and biosecurity experts to find a universal jailbreak that defeats its five-question bio safety challenge. The model in scope is GPT-5.5 in Codex Desktop only, with a $25,000 reward for the first true universal jailbreak and smaller awards possible for partial wins. Applications run from April 23 to June 22, 2026, testing runs from April 28 to July 27, 2026, and all findings are covered by NDA.
OpenAI BlogPrompt injection turned Google’s Antigravity file search into RCE
Apr 21, 2026MediumNewsSecurityIndustryPillar Security researchers disclosed a prompt injection flaw in Google's Antigravity IDE that can turn the find_my_name file search tool into remote code execution. The Pattern parameter accepts strings beginning with "-", which the fd utility reads as flags, and Secure Mode never evaluates the call because it runs before the security boundary.
Fix: Google has already fixed the flaw internally, and Antigravity users need not do anything else to remain protected.
CSO OnlineGoogle Patches Antigravity IDE Flaw Enabling Prompt Injection Code Execution
Apr 21, 2026MediumNewsSecurityIndustryCybersecurity researchers at Pillar Security found a flaw in Google's agentic IDE Antigravity that allows code execution. The find_by_name file-search tool passes its Pattern parameter directly to fd without strict validation, so injecting the -X (exec-batch) flag, such as the value -Xsh, makes fd run matched files as shell scripts, bypassing Strict Mode. An attacker can stage a malicious file and trigger it, or reach the same result through indirect prompt injection from an untrusted file.
Fix: Fixed by Google; the flaw was reported on January 7, 2026 and addressed as of February 28.
The Hacker NewsCursor AI Vulnerability Exposed Developer Devices
Apr 17, 2026MediumNewsSecurityIndustryResearchers reported that an indirect prompt injection could be chained with a sandbox bypass and Cursor's remote tunnel feature to gain shell access to developer machines.
SecurityWeekClaude Code, Gemini CLI, GitHub Copilot Agents Vulnerable to Prompt Injection via Comments
Apr 16, 2026MediumNewsSecuritySafetyA researcher has disclosed details of an AI attack method he calls 'Comment and Control'. The method reportedly enables prompt injection through comments in Claude Code, Gemini CLI and GitHub Copilot Agents.
SecurityWeekCopilot and Agentforce fall to form-based prompt injection tricks
Apr 15, 2026MediumNewsSecurityPrivacyCapsule Security researchers disclosed prompt-injection flaws in Microsoft Copilot Studio and Salesforce Agentforce that let crafted input override agent instructions and exfiltrate data. In the Microsoft case, a payload in a SharePoint form field leaks customer records from connected SharePoint Lists by email, and the flaw was assigned CVE-2026-21520 with a CVSS score of 7.5. In the Salesforce case, a malicious public lead form causes an agent to pull CRM records via the "GetLeadsInformation" function and email them externally, and Salesforce called the vector "configuration-specific" and pointed to optional human-in-the-loop controls.
Fix: Microsoft patched the issue following disclosure; the mitigation was carried out internally and no further action is required from users. For the broader issue, the source says both disclosures converge on treating all external inputs as untrusted, using filters that separate data from instructions, enforcing input validation, least-privilege access, and strict controls on actions like outbound email.
CSO OnlineZero‑click Grafana AI attack can enable enterprise data exfiltration
Apr 7, 2026MediumNewsSecurityPrivacyNoma Security disclosed GrafanaGhost, a chained exploit in Grafana's AI components that can leak sensitive data such as financial metrics, infrastructure health data, customer data and operational logs without credentials or user interaction. The chain combines an indirect prompt injection with a URL validation flaw that lets protocol-relative URLs like //attacker.com bypass client-side image-loading protections, and it uses the keyword INTENT to evade AI guardrails. Grafana reportedly validated the flaw and rolled out a fix, though it did not immediately respond to CSO's request for comment.
Fix: Fix rolled out by Grafana (version not stated). BeyondTrust's Bradley Smith recommends checking whether Grafana AI/LLM features are enabled, patching to the latest version, restricting "img-src" to known domains, and applying egress controls.
CSO OnlineGoogle Workspace’s continuous approach to mitigating indirect prompt injections
Apr 2, 2026InfoNewsSecurityResearchGoogle's GenAI Security Team describes its continuous approach to defending Workspace with Gemini against indirect prompt injection (IPI), where malicious instructions placed in data or tools steer an LLM during a user's query. The post covers discovery methods, including human and automated red-teaming, a Vulnerability Rewards Program and monitoring of public disclosures, plus a catalog process for new vulnerabilities. It also reports that its Simula synthetic data generation boosted output by 75%.
Google Online Security BlogClaude Extension Flaw Enabled Zero-Click XSS Prompt Injection via Any Website
Mar 26, 2026MediumNewsSecuritySafetyKoi Security researcher Oren Yomtov disclosed ShadowPrompt, a flaw in Anthropic's Claude Google Chrome Extension that lets any website silently inject prompts into the assistant. The issue chains an overly permissive origin allowlist matching *.claude.ai with a DOM-based XSS flaw in an Arkose Labs CAPTCHA component hosted on a-cdn.claude[.]ai. A successful attack could steal access tokens, expose conversation history, and perform actions as the victim, with no clicks or permission prompts required.
Fix: Anthropic deployed a patch to the Chrome extension (version 1.0.41) that enforces a strict origin check requiring an exact match to the domain claude[.]ai. Arkose Labs fixed the XSS flaw as of February 19, 2026.
The Hacker NewsOpenClaw AI Agent Flaws Could Enable Prompt Injection and Data Exfiltration
Mar 14, 2026MediumNewsSecuritySafetyChina's CNCERT warned that OpenClaw, an open-source self-hosted autonomous AI agent, has weak default security configurations and privileged system access that attackers could exploit to take control of endpoints. Indirect prompt injection can trick the agent into leaking data, and PromptArmor showed that link previews in messaging apps like Telegram or Discord can carry data to an attacker's domain without a click.
Fix: CNCERT advises users and organizations to strengthen network controls, prevent exposure of OpenClaw's default management port to the internet, isolate the service in a container, avoid storing credentials in plaintext, download skills only from trusted channels, disable automatic updates for skills, and keep the agent up-to-date.
The Hacker News
Topic added 2026-10-09. An item belongs to this topic when its title matches one of the topic's patterns or its summary mentions the topic at least twice. Report a wrong match with the feedback button on the item.