Prompt injection and jailbreaks
Inputs that override a model's instructions, directly or through content it reads, and attempts to bypass its safeguards.
- All items
- 192
- Last 90 days
- 41
- Change
- -21%vs 52 before
Items per month
| Month | Items |
|---|---|
| May 2025 | 2 |
| Jun 2025 | 0 |
| Jul 2025 | 3 |
| Aug 2025 | 24 |
| Sep 2025 | 0 |
| Oct 2025 | 6 |
| Nov 2025 | 3 |
| Dec 2025 | 1 |
| Jan 2026 | 2 |
| Feb 2026 | 7 |
| Mar 2026 | 10 |
| Apr 2026 | 19 |
| May 2026 | 7 |
| Jun 2026 | 15 |
| Jul 2026 | 26 |
| Aug 2026 | 18 |
| Sep 2026 | 7 |
| Oct 2026 | 6 |
192 items
New Gaslight macOS Malware Uses Prompt Injection to Disrupt AI-Assisted Analysis
Jun 25, 2026MediumNewsSecuritySafetySentinelOne researchers report Gaslight, a previously undocumented Rust-based macOS implant and information stealer that they assess with high confidence as the work of North Korea-aligned threat actors. Its embedded prompt injection payload contains 38 fabricated system messages meant to make LLM-assisted malware triage agents abort, truncate, or refuse analysis. The malware also collects Terminal histories, Keychain data and browser data, compresses it into a ZIP archive, and uploads it over a Telegram bot C2 channel.
The Hacker NewsPrompt Injection as Role Confusion
Jun 22, 2026LowNewsSecurityResearchCharles Ye, Jasmine Cui, and Dylan Hadfield-Menell study whether models can tell privileged role-tagged text, such as <system>, <think> and <assistant>, from untrusted <user> input. They report that models weigh the writing style of text more heavily than its actual content, enabling jailbreaks like the one against gpt-oss-20b. Rewriting text to be less stylistically similar to the expected role format cut average attack success from 61% to 10%.
Simon Willison's WeblogSecurity considerations for adopting Claude Code and Cowork for SMBs
Jun 19, 2026InfoNewsSecurityIndustryA security leader at a small or medium-sized business shares lessons on adopting Claude, Code and Cowork. The piece advises matching the Claude plan to real needs, since key security controls such as the Compliance API appear only on the Enterprise plan, and it recommends phased feature enablement, noting that indirect prompt injection risk from features like web search and browser extensions is real and still emerging.
CSO OnlineM365 Copilot SearchLeak: Your prompt injection attack surface just got bigger
Jun 19, 2026MediumNewsSecurityIndustryVaronis Threat Labs researchers disclosed SearchLeak, a proof-of-concept attack against Microsoft's M365 Copilot Enterprise Search that leaks sensitive corporate data when employees click crafted links. The attack chains three weaknesses, using the ?q= URL parameter as a natural-language prompt that can instruct the LLM to surface and exfiltrate accessible business content such as emails, SharePoint and OneDrive files. Microsoft rated the information disclosure flaw as critical and patched it server-side.
Fix: Microsoft patched the vulnerability on the server side earlier this month.
CSO OnlineThe US government’s Anthropic models ban was never about an AI jailbreak
Jun 15, 2026InfoNewsPolicyIndustryOn Friday, the U.S. Commerce Department sent Anthropic a letter invoking an export control directive that barred non-Americans, including Anthropic's employees, from accessing Fable 5 and Mythos 5, citing an unspecified national security concern. Anthropic shut both models down for all customers to comply, and said it believes the letter relates to a guardrail bypass in Fable 5 but cannot confirm this because the letter gives no details. Security researcher Katie Moussouris argued the bypass should never have triggered an export control.
TechCrunch (Security)5 runtime signals for catching a compromised AI agent
Jun 15, 2026LowNewsSecuritySafetySimon Willison described the lethal trifecta, three capabilities that combined in one AI agent (access to private data, exposure to untrusted content, and the ability to communicate externally) create a near-guaranteed path to exploitation through indirect prompt injection. The article argues these capabilities are now the default configuration for useful agents rather than edge cases, so the trifecta is no longer a meaningful risk indicator on its own.
Fix: Meta's "Rule of Two" framework, published in October 2025, recommends that agents satisfy no more than two of the three trifecta properties in a single session, with human-in-the-loop approval required if all three are necessary. The source notes Meta concedes this framework may not cleanly fit many use cases and that designs satisfying it can still fail. McKerchar's "blast radius reduction" is also mentioned as an operational philosophy.
CSO OnlinePrompt injection breaks today’s AI agents, study warns
Jun 12, 2026LowNewsSecurityResearchResearchers from Nanyang Technological University, ST Engineering, IBM Research and the University of Illinois Urbana-Champaign released StakeBench, a benchmark that tested prompt injection against web agents NanoBrowser and BrowserUse across 3,168 adversarial runs. Indirect injection hidden in web content succeeded 41.67% to 68.16% of the time, and direct injection exceeded 79% across all configurations, with no scenario consistently blocked in GPT-5 and Gemini systems.
CSO OnlineAnthropic Disputes Fable 5 AI Jailbreak
Jun 12, 2026LowNewsSafetySecurityAnthropic disputed claims by the researcher Pliny the Liberator that they jailbroke Claude Fable 5, a model launched as generally available. The researcher said multi-agent prompting elicited information on sensitive topics and published screenshots and the allegedly internal system prompt. Anthropic said the demonstrated method only coaxes the model past its refusals, that independent classifier systems stay in force, and that some outputs were not from Fable 5 while the rest offered only public general information.
SecurityWeekOpenAI unveils Lockdown Mode to protect sensitive data from prompt injection attacks
Jun 6, 2026InfoNewsSecuritySafetyOpenAI announced Lockdown Mode, a ChatGPT feature meant to protect sensitive data from prompt injection attacks hidden in webpages and other content. It disables live web browsing, retrieval and display of web images, deep research, and agent mode. OpenAI says ChatGPT could still be affected by prompt injections in cached web content or uploaded files, and the feature is rolling out to self-serve ChatGPT Business accounts and eligible personal accounts.
Fix: Lockdown Mode, which OpenAI says is designed for people and organizations that handle sensitive data and want stricter protection from data exfiltration risks related to prompt injection.
TechCrunch (Security)OpenAI Help: Lockdown Mode
Jun 5, 2026InfoNewsSecuritySafetyOpenAI has released Lockdown Mode for ChatGPT, rolling out to eligible personal accounts (Free, Go, Plus, Pro) and self-serve Business accounts. It is designed to block the final stage of data exfiltration from a prompt injection attack by limiting outbound network requests that could send sensitive data to an attacker. It does not stop prompt injections from appearing in processed content, such as cached web content or uploaded files, which can still affect response behavior or accuracy.
Fix: Lockdown Mode, which limits outbound network requests that could transfer sensitive data to an attacker.
Simon Willison's WeblogAmazon Q Developer and Kiro – Prompt Injection Issues in Kiro and Q IDE plugins
Jun 5, 2026HighVulnerabilitySecuritySafetyAWS published bulletin AWS-2025-019 on prompt injection issues in Amazon Q Developer and Kiro, reported in Embrace The Red's "The Month of AI Bugs" blog posts. The Amazon Q Developer flaws, affecting find, grep and echo commands (versions below 1.22.0) and ping and dig commands (versions below 1.24.0), could run commands or exfiltrate metadata via DNS without Human-in-the-Loop confirmation. The Kiro flaw requires local system access and can lead to arbitrary code execution via IDE or MCP settings files, in Autopilot or Supervised mode, affecting version 0.1.42.
Fix: Amazon Q Developer Language Server v1.22.0 (released July 17, 2025) requires HITL confirmation for find, grep and echo commands. Language Server v1.24.0 (released July 29, 2025) requires HITL confirmation for ping and dig commands. Kiro version 0.1.42 (released August 1, 2025) requires HITL confirmation for these actions when configured in Supervised mode. AWS also recommends that customers evaluate and implement appropriate security controls and policies for their environments.
AWS Security BulletinsClaude Code GitHub Action Flaw Let One Malicious Issue Hijack Repositories
Jun 4, 2026MediumNewsSecurityIndustryA security researcher found that Anthropic's Claude Code GitHub Action let anyone running a GitHub App bypass the write-access check, so a single opened issue could trigger the action on public repositories. Through indirect prompt injection, the attacker could make Claude write /proc/self/environ values, including the GitHub Actions OIDC credentials, back into the issue, enabling write access to the target's code, issues and workflows. Anthropic fixed the core bypass within four days of a January report and rated the issues 7.8 under CVSS v4.0.
Fix: Update to claude-code-action v1.0.94 or later. Then audit any workflow that lets users without write access, or bots, trigger Claude: if it takes untrusted input, don't feed it any secret beyond the Anthropic API key and GITHUB_TOKEN, and remove tools and permissions that can be used for exfiltration.
The Hacker NewsShadow AI: The Hidden Risk Expanding Across the Enterprise
May 29, 2026InfoNewsSecurityIndustryEmployees and developers are adopting GenAI tools and AI features faster than security teams can review them, creating shadow AI that can leak sensitive data. The source says indirect prompt injection is especially dangerous because malicious instructions can be hidden in documents, websites or knowledge bases. It argues traditional tools like web proxies, firewalls and Zero Trust do not cover AI interactions.
Fix: The source promotes CrowdStrike Falcon® AI Detection and Response (AIDR) for visibility, control and protection against AI-specific threats such as prompt injection. It does not describe specific configuration steps or fixes.
CrowdStrike BlogNew image-based prompt injection attack targets multimodal AI models
May 18, 2026InfoNewsSecurityResearchXidian University researchers published a paper describing CrossMPI, an image-only prompt injection technique that uses nearly imperceptible pixel perturbations to alter how large vision-language models (LVLMs) interpret both visual and textual inputs without changing the text prompt. In tests against open-source LVLMs including MiniGPT4, BLIP-2, InstructBLIP, BLIVA and Qwen2.5-VL, the attack reportedly achieved an average success rate of 66.36%, roughly 41 percentage points above prior baselines. The paper claims strong black-box transferability.
CSO OnlineGHSA-72w5-pf8h-xfp4: DeepSeek TUI: task_create Insecure Defaults Enable RCE via Prompt Injection in Project Files
May 14, 2026CriticalVulnerabilitySecurityCVE-2026-45374The task_create tool in DeepSeek TUI spawns durable sub-agents that default to allow_shell=true (config.rs:1499) and auto_approve=true (task_manager.rs:297). A user who approves a benign-looking task_create call therefore unknowingly grants the sub-agent unapproved shell access, which can follow attacker-controlled instructions in a cloned repository's AGENTS.md file and run commands without a further approval prompt. The reporter demonstrates remote code execution through a proof of concept that triggers a callback to a collaborator server.
Fix: Default allow_shell to false for durable tasks (config.rs:1499: self.allow_shell.unwrap_or(false)); default auto_approve to false for durable tasks (task_manager.rs:297: auto_approve: None, instead of Some(true)); and, when the model requests task_create with allow_shell=true, surface that in the approval prompt so the user knows they are granting shell access.
GitHub Advisory DatabaseWhen prompts become shells: RCE vulnerabilities in AI agent frameworks
May 7, 2026MediumNewsSecurityResearchMicrosoft's Semantic Kernel, an open-source framework for building AI agents, contained two vulnerabilities, CVE-2026-25592 and CVE-2026-26030, which the source says have since been fixed. The flaws let an attacker use prompt injection to reach unauthorized code execution on the host running an agent, and the source demonstrates this by launching calc.exe with a single prompt. Exploitation of CVE-2026-26030 requires a prompt injection vector and an agent with the Search Plugin backed by the In-Memory Vector Store.
Fix: Fixed in Semantic Kernel (the source states the flaws "have since been fixed" but does not give fixed version numbers). The source also says customers should assess exposure, patch affected agents, and investigate whether exploitation may already have occurred, but it does not detail those steps.
Microsoft Security BlogPoisoned truth: The quiet security threat inside enterprise AI
May 6, 2026InfoNewsSecurityResearchEnterprise AI security discussions often focus on prompt injection and jailbreaks, but SANS Institute leaders argue that corrupted model understanding, known as AI data poisoning, deserves more attention. The term covers deliberate manipulation of training data, RAG pipelines, knowledge bases or agent memory, as well as stale or conflicting enterprise data. Chris Cochran of SANS says the damage is hard to spot because the business appears to operate normally while the system produces plausible but wrong answers.
CSO OnlineVLM-Guard: Defending Jailbreaks by Monitoring Only Hundreds of Safety-Critical Neurons
May 4, 2026InfoResearchPeer-reviewedSecurityResearchVLM-Guard is a jailbreak detection framework for Large Vision Language Models that identifies safety-critical neurons linked to harmful semantics through differential analysis of activation values. It isolates a compact set of just a few hundred neurons, less than 0.2% of the total, to build a detector that is lightweight and training-free. The authors report that it detects jailbreak attacks effectively while preserving benign performance in attack-free settings.
IEEE Xplore (Security & AI Journals)Hugging Face, ClawHub Abused for Malware Distribution
May 1, 2026MediumNewsSecurityIndustryThreat actors are distributing malware through trojanized shared files on Hugging Face and ClawHub, according to Acronis. The attacks rely on social engineering to get users to download files that execute commands, fetch payloads and install hidden dependencies. Acronis identified close to 600 malicious skills across 13 ClawHub developer accounts, with the OpenClaw ecosystem's indirect prompt injection used to make agents download and run code.
SecurityWeekMalicious AI Prompt Injection Attacks Increasing, but Sophistication Still Low: Google
Apr 27, 2026InfoNewsSecurityResearchGoogle researchers scanned Common Crawl website snapshots for known indirect prompt injection patterns and used Gemini and human review to remove false positives. They found pranks, attempts to deter AI crawlers, SEO manipulation, helpful guidance, and some malicious attacks, mostly low in sophistication. The malicious attempts were exfiltration of data such as IPs and credentials to an attacker-specified email address, and destruction prompts aimed at deleting files, which the researchers considered unlikely to succeed. Malicious attempts rose 32% between November 2025 and February 2026, and the researchers expect scale and complexity to grow.
SecurityWeek
Topic added 2026-10-09. An item belongs to this topic when its title matches one of the topic's patterns or its summary mentions the topic at least twice. Report a wrong match with the feedback button on the item.