A fundamental flaw leaves LLMs strikingly vulnerable to attack
Summary
Researchers discovered a fundamental flaw in how large language models (LLMs, AI systems trained on text to generate responses) identify the source of instructions, making them vulnerable to chain-of-thought forgery attacks (tricking an LLM by mimicking the internal notes it writes to itself). By exploiting this flaw, attackers can trick popular LLMs into providing dangerous information they were trained to refuse, such as instructions for making drugs or sabotaging aircraft, and the researchers argue this vulnerability may be fundamentally impossible to fully secure against.
Classification
Affected Vendors
Related Issues
Original source: https://www.technologyreview.com/2026/07/30/1140927/a-fundamental-flaw-leaves-llms-vulnerable-to-attack/
First tracked: July 30, 2026 at 08:01 AM
Classified by LLM (prompt v3) · confidence: 92%