{"data":{"id":"e6515a3e-cc27-4f09-a455-903021ab6a68","title":"Stealing Reasoning Traces from Proprietary LLM APIs","summary":"Researchers discovered that AI companies like OpenAI, Anthropic, and Google were returning encrypted reasoning traces (the step-by-step thinking process an AI uses to solve problems) to users in a way that could be replayed and exploited. By feeding these encrypted blocks into weaker versions of the same AI models with jailbreak prompts (tricking the AI into ignoring safety guidelines), attackers could extract the stronger model's hidden reasoning in readable form, and even use a prompt injection attack (hiding malicious instructions within the reasoning traces) to make models perform unintended actions like exfiltrating data.","solution":"All model providers acknowledged the report and subsequently the vulnerability was unable to be reproduced in follow-up testing, indicating the issue has been fixed. Specifically, the jailbreak technique that worked on Claude Haiku 4.5 (using a prompt to transcribe reasoning verbatim) no longer works in Haiku 4.6 models, as that feature was removed.","labels":["security","research"],"sourceUrl":"https://simonwillison.net/2026/Aug/11/stealing-reasoning-traces/","publishedAt":"2026-08-11T22:40:45.000Z","cveId":null,"cweIds":null,"cvssScore":null,"cvssSeverity":null,"severity":"high","attackType":["model_theft","data_extraction","prompt_injection"],"issueType":"news","affectedPackages":null,"affectedVendors":["OpenAI","Anthropic","Google"],"affectedVendorsRaw":["OpenAI","Anthropic","Google","GPT-5.5","Claude Haiku 4.5"],"classifierModel":"claude-haiku-4-5-20251001","classifierPromptVersion":"v3","cvssVector":null,"attackVector":null,"attackComplexity":null,"privilegesRequired":null,"userInteraction":null,"exploitMaturity":null,"epssScore":null,"patchAvailable":null,"disclosureDate":"2026-08-11T22:40:45.000Z","capecIds":null,"crossRefCount":0,"attackSophistication":"advanced","impactType":["confidentiality","integrity"],"aiComponentTargeted":"api","llmSpecific":true,"classifierConfidence":0.95,"researchCategory":null,"atlasIds":null}}