Encrypted Prompts Bypass AI Safety Guardrails in Grok and Gemini
Summary
Researchers discovered cryptographic context injection, an attack where encrypted prompts bypass safety guardrails (automated systems that block harmful requests) in AI models like Grok and Gemini. The attack works by hiding malicious instructions inside encrypted text, which safety filters cannot read, then decrypting it inside the model's code execution sandbox (a contained environment where code runs safely), allowing the AI to follow harmful instructions it would normally refuse. The attack can be delivered directly to chat or indirectly through weaponized web pages that trick AI agents into processing the encrypted payload.
Solution / Mitigation
Adversa's report includes prevention advice for defenders, but the source text does not explicitly describe or quote any specific mitigation steps, fixes, or updates.
Classification
Affected Vendors
Related Issues
CVE-2024-27444: langchain_experimental (aka LangChain Experimental) in LangChain before 0.1.8 allows an attacker to bypass the CVE-2023-
CVE-2026-30308: In its design for automatic terminal command execution, HAI Build Code Generator offers two options: Execute safe comman
Original source: https://www.securityweek.com/encrypted-prompts-bypass-ai-safety-guardrails-in-grok-and-gemini/
First tracked: August 21, 2026 at 02:00 PM
Classified by LLM (prompt v3) · confidence: 92%