Prompt injection and jailbreaks
Inputs that override a model's instructions, directly or through content it reads, and attempts to bypass its safeguards.
- All items
- 194
- Last 90 days
- 43
- Change
- -17%vs 52 before
Items per month
| Month | Items |
|---|---|
| May 2025 | 2 |
| Jun 2025 | 0 |
| Jul 2025 | 3 |
| Aug 2025 | 24 |
| Sep 2025 | 0 |
| Oct 2025 | 6 |
| Nov 2025 | 3 |
| Dec 2025 | 1 |
| Jan 2026 | 2 |
| Feb 2026 | 7 |
| Mar 2026 | 10 |
| Apr 2026 | 19 |
| May 2026 | 7 |
| Jun 2026 | 15 |
| Jul 2026 | 26 |
| Aug 2026 | 18 |
| Sep 2026 | 7 |
| Oct 2026 | 8 |
194 items
Hugging Face, ClawHub Abused for Malware Distribution
May 1, 2026MediumNewsSecurityIndustryThreat actors are distributing malware through trojanized shared files on Hugging Face and ClawHub, according to Acronis. The attacks rely on social engineering to get users to download files that execute commands, fetch payloads and install hidden dependencies. Acronis identified close to 600 malicious skills across 13 ClawHub developer accounts, with the OpenClaw ecosystem's indirect prompt injection used to make agents download and run code.
SecurityWeekMalicious AI Prompt Injection Attacks Increasing, but Sophistication Still Low: Google
Apr 27, 2026InfoNewsSecurityResearchGoogle researchers scanned Common Crawl website snapshots for known indirect prompt injection patterns and used Gemini and human review to remove false positives. They found pranks, attempts to deter AI crawlers, SEO manipulation, helpful guidance, and some malicious attacks, mostly low in sophistication. The malicious attempts were exfiltration of data such as IPs and credentials to an attacker-specified email address, and destruction prompts aimed at deleting files, which the researchers considered unlikely to succeed. Malicious attempts rose 32% between November 2025 and February 2026, and the researchers expect scale and complexity to grow.
SecurityWeekGHSA-rp7v-4384-hfrp: k8sGPT has Prompt Injection through its k8sGPT-Operator
Apr 24, 2026HighVulnerabilitySecuritySafetyk8sGPT's auto-remediation pipeline deserializes AI-generated YAML directly into a Deployment object in `object_to_execution.go`, without validating it against the original Deployment object. The issue was fixed after coordination with Alex Jones, and the proof of concept was shared only with the maintainers.
GitHub Advisory DatabaseCVE-2026-41318: AnythingLLM stored XSS through chart captions in chat history
Apr 24, 2026MediumVulnerabilitySecurityCVE-2026-41318AnythingLLM versions prior to 1.12.1 contain a stored DOM-level XSS flaw, tracked as CVE-2026-41318. The in-chat markdown renderer interpolates image `alt` text into an HTML attribute without encoding, and the `Chartable` component renders chart captions with no DOMPurify sanitization. An attacker who can influence LLM output, through indirect prompt injection in a shared workspace document or by creating a chart record in a multi-user workspace, can run script in other users' browsers when they open that conversation.
Fix: Patched in version 1.12.1.
NVD/CVE DatabaseAI threats in the wild: The current state of prompt injections on the web
Apr 23, 2026MediumNewsSecurityResearchGoogle's Threat Intelligence teams swept Common Crawl, a public repository of mostly static English-language websites, to check whether real attackers are using indirect prompt injection (IPI) on the open web. The source describes a coarse-to-fine filtering approach, starting with pattern matching on phrases such as "ignore … instructions," because naive searches return mostly benign research and educational content. The article is cut off before reporting its findings.
Google Online Security BlogGPT-5.5 Bio Bug Bounty
Apr 22, 2026InfoNewsSafetyIndustryOpenAI is launching a Bio Bug Bounty for GPT-5.5, inviting AI red teamers, security researchers and biosecurity experts to find a universal jailbreak that defeats its five-question bio safety challenge. The model in scope is GPT-5.5 in Codex Desktop only, with a $25,000 reward for the first true universal jailbreak and smaller awards possible for partial wins. Applications run from April 23 to June 22, 2026, testing runs from April 28 to July 27, 2026, and all findings are covered by NDA.
OpenAI BlogGHSA-2r2p-4cgf-hv7h: engram: HTTP server CORS wildcard + auth-off-by-default enables CSRF graph exfiltration and persistent indirect prompt injection
Apr 22, 2026HighVulnerabilitySecuritySafetyThe local HTTP server started by `engram server` (default `127.0.0.1:7337`) accepted requests from any browser origin with no authentication unless `ENGRAM_API_TOKEN` was set. Combined with `Access-Control-Allow-Origin: *` and a body parser that did not require `Content-Type: application/json`, a malicious web page could read the knowledge graph via `GET /query` and `GET /stats`, and write persistent prompt-injection payloads via `POST /learn`, which were later surfaced to the user's AI coding agent. Affected versions are `engramx` >= 1.0.0 and < 2.0.2.
Fix: Fixed in `engramx@2.0.2`. Remediation in 2.0.2 includes fail-closed auth on non-public routes (Bearer header or HttpOnly cookie, constant-time comparison, 256-bit token at `~/.engram/http-server.token`), removal of wildcard CORS with an opt-in allowlist via `ENGRAM_ALLOWED_ORIGINS`, Host and Origin validation, enforced `Content-Type: application/json` on mutations, and a `/ui?token=` bootstrap with a `Sec-Fetch-Site` gate. Workarounds if upgrading is not possible: do not run `engram server` or `engram ui`, or set `ENGRAM_API_TOKEN` to a long random value and stop the server before browsing the web.
GitHub Advisory DatabaseGHSA-3hjv-c53m-58jj: Flowise: CSV Agent Prompt Injection Remote Code Execution Vulnerability
Apr 21, 2026CriticalVulnerabilitySecurityCVE-2026-41264Trend Micro's Zero Day Initiative reports an unauthenticated remote code execution flaw in FlowiseAI Flowise, tested at version 3.0.13. The flaw sits in the run method of the CSV_Agents class, which evaluates an LLM-generated Python script without proper sandboxing. The input validation that checks for forbidden patterns can be bypassed, allowing arbitrary OS commands to run on the server as the user running it.
GitHub Advisory DatabasePrompt injection turned Google’s Antigravity file search into RCE
Apr 21, 2026MediumNewsSecurityIndustryPillar Security researchers disclosed a prompt injection flaw in Google's Antigravity IDE that can turn the find_my_name file search tool into remote code execution. The Pattern parameter accepts strings beginning with "-", which the fd utility reads as flags, and Secure Mode never evaluates the call because it runs before the security boundary.
Fix: Google has already fixed the flaw internally, and Antigravity users need not do anything else to remain protected.
CSO OnlineGoogle Patches Antigravity IDE Flaw Enabling Prompt Injection Code Execution
Apr 21, 2026MediumNewsSecurityIndustryCybersecurity researchers at Pillar Security found a flaw in Google's agentic IDE Antigravity that allows code execution. The find_by_name file-search tool passes its Pattern parameter directly to fd without strict validation, so injecting the -X (exec-batch) flag, such as the value -Xsh, makes fd run matched files as shell scripts, bypassing Strict Mode. An attacker can stage a malicious file and trigger it, or reach the same result through indirect prompt injection from an untrusted file.
Fix: Fixed by Google; the flaw was reported on January 7, 2026 and addressed as of February 28.
The Hacker NewsCursor AI Vulnerability Exposed Developer Devices
Apr 17, 2026MediumNewsSecurityIndustryResearchers reported that an indirect prompt injection could be chained with a sandbox bypass and Cursor's remote tunnel feature to gain shell access to developer machines.
SecurityWeekGHSA-6r77-hqx7-7vw8: Flowise: APIChain Prompt Injection SSRF in GET/POST API Chains
Apr 16, 2026HighVulnerabilitySecurityFlowiseAI's POST and GET API Chain components build request URLs from LLM output without validating them against the intended API documentation. Unauthenticated attackers can inject a crafted documentation prompt that overrides the BASE URL, causing the server to send arbitrary HTTP requests to internal and external hosts, as demonstrated against /flag on host.docker.internal:8080. The source affects FlowiseAI instances at version 2.2.1 and below.
GitHub Advisory DatabaseClaude Code, Gemini CLI, GitHub Copilot Agents Vulnerable to Prompt Injection via Comments
Apr 16, 2026MediumNewsSecuritySafetyA researcher has disclosed details of an AI attack method he calls 'Comment and Control'. The method reportedly enables prompt injection through comments in Claude Code, Gemini CLI and GitHub Copilot Agents.
SecurityWeekCVE-2026-30615: Windsurf prompt injection allows arbitrary command execution via HTML content
Apr 15, 2026CriticalVulnerabilitySecurityCVE-2026-30615A prompt injection flaw in Windsurf 1.9544.26 lets remote attackers run arbitrary commands on a victim system. When Windsurf processes attacker-controlled HTML, injected instructions can modify the local MCP configuration and automatically register a malicious MCP STDIO server, with no further user interaction. Successful exploitation can execute commands as the user, persist the malicious configuration, and expose sensitive information accessible through the application.
NVD/CVE DatabaseCopilot and Agentforce fall to form-based prompt injection tricks
Apr 15, 2026MediumNewsSecurityPrivacyCapsule Security researchers disclosed prompt-injection flaws in Microsoft Copilot Studio and Salesforce Agentforce that let crafted input override agent instructions and exfiltrate data. In the Microsoft case, a payload in a SharePoint form field leaks customer records from connected SharePoint Lists by email, and the flaw was assigned CVE-2026-21520 with a CVSS score of 7.5. In the Salesforce case, a malicious public lead form causes an agent to pull CRM records via the "GetLeadsInformation" function and email them externally, and Salesforce called the vector "configuration-specific" and pointed to optional human-in-the-loop controls.
Fix: Microsoft patched the issue following disclosure; the mitigation was carried out internally and no further action is required from users. For the broader issue, the source says both disclosures converge on treating all external inputs as untrusted, using filters that separate data from instructions, enforcing input validation, least-privilege access, and strict controls on actions like outbound email.
CSO OnlineZero‑click Grafana AI attack can enable enterprise data exfiltration
Apr 7, 2026MediumNewsSecurityPrivacyNoma Security disclosed GrafanaGhost, a chained exploit in Grafana's AI components that can leak sensitive data such as financial metrics, infrastructure health data, customer data and operational logs without credentials or user interaction. The chain combines an indirect prompt injection with a URL validation flaw that lets protocol-relative URLs like //attacker.com bypass client-side image-loading protections, and it uses the keyword INTENT to evade AI guardrails. Grafana reportedly validated the flaw and rolled out a fix, though it did not immediately respond to CSO's request for comment.
Fix: Fix rolled out by Grafana (version not stated). BeyondTrust's Bradley Smith recommends checking whether Grafana AI/LLM features are enabled, patching to the latest version, restricting "img-src" to known domains, and applying egress controls.
CSO OnlineGoogle Workspace’s continuous approach to mitigating indirect prompt injections
Apr 2, 2026InfoNewsSecurityResearchGoogle's GenAI Security Team describes its continuous approach to defending Workspace with Gemini against indirect prompt injection (IPI), where malicious instructions placed in data or tools steer an LLM during a user's query. The post covers discovery methods, including human and automated red-teaming, a Vulnerability Rewards Program and monitoring of public disclosures, plus a catalog process for new vulnerabilities. It also reports that its Simula synthetic data generation boosted output by 75%.
Google Online Security BlogGHSA-44c2-3rw4-5gvh: PraisonAI Has SSRF in FileTools.download_file() via Unvalidated URL
Apr 1, 2026HighVulnerabilitySecurityCVE-2026-34954GHSA-44c2-3rw4-5gvh reports an SSRF flaw in FileTools.download_file() in praisonaiagents. The function checks the destination path but passes the caller-supplied url directly to httpx.stream() with follow_redirects=True, so an attacker who controls the URL can reach any host the server can access, including cloud metadata services and internal services. On EC2 instances with IMDSv1 enabled, IAM credentials can be retrieved and written to disk, and the flaw is reachable through indirect prompt injection without authentication.
Fix: Suggested fix: validate the url before the request by allowing only the http and https schemes and blocking loopback, link-local (169.254.0.0/16), and private RFC 1918 address ranges, as shown in the source's _validate_url() example.
GitHub Advisory DatabaseGHSA-w37c-qqfp-c67f: PraisonAI: Shell Injection in run_python() via Unescaped $() Substitution
Apr 1, 2026HighVulnerabilitySecurityCVE-2026-34937PraisonAI's `run_python()` in `praisonai` builds a shell command by placing user-controlled code inside `python3 -c "<code>"` and runs it with `subprocess.run(..., shell=True)`. Its escaping only handles `\` and `"`, so `$()` and backtick substitutions execute as OS commands before Python starts, enabling arbitrary command execution as the process user. The advisory notes the function is reachable through indirect prompt injection, and the auto-generated Flask server ships with `AUTH_ENABLED = False` when no token is configured.
GitHub Advisory DatabaseGHSA-6vh2-h83c-9294: PraisonAI: Python Sandbox Escape via str Subclass startswith() Override in execute_code
Apr 1, 2026CriticalVulnerabilitySecurityCVE-2026-34938The execute_code() function in praisonai-agents runs attacker-controlled Python inside a three-layer sandbox. Passing a str subclass with an overridden startswith() method to the _safe_getattr wrapper bypasses the sandbox, allowing arbitrary OS command execution on the host as the process user. Deployments using bot.py, autonomy_mode.py, or bots_cli.py set PRAISONAI_AUTO_APPROVE=true by default, so the tool can fire without human confirmation when triggered via indirect prompt injection.
GitHub Advisory Database
Topic added 2026-10-09. An item belongs to this topic when its title matches one of the topic's patterns or its summary mentions the topic at least twice. Report a wrong match with the feedback button on the item.