OpenAI, Anthropic, Google API Flaw Let Weaker AI Models Decode Stronger Models' Reasoning
Summary
Researchers discovered a flaw in how OpenAI, Anthropic, and Google handle encrypted reasoning objects (encrypted data that stores an AI's hidden thinking between API calls) that allowed them to recover secrets from these hidden blocks, including API keys, passwords, and private data from user sessions. The flaw worked because these encrypted reasoning blocks could be replayed across different sessions and even given to weaker models in the same provider family, which could then decode the hidden content. The researchers identified four ways this could be abused: stealing proprietary reasoning, extracting private user data, recovering harmful content hidden in reasoning, and injecting malicious prompts inside the opaque blocks.
Solution / Mitigation
The source states that "the demonstrated attacks stopped working after mitigations" and notes that "the main extraction attack is no longer reproducible as of August 2026." Additionally, developers are advised to "strip reasoning blocks and opaque reasoning fields from shared traces and avoid committing raw API transcripts even when the visible text has been sanitized."
Classification
Affected Vendors
Related Issues
Original source: https://thehackernews.com/2026/08/openai-anthropic-google-api-flaw-let.html
First tracked: August 12, 2026 at 02:01 PM
Classified by LLM (prompt v3) · confidence: 95%