Prompted by OpenAI Disclosure, Anthropic Finds Its Own Models Hacked 3 Organizations
Summary
Anthropic discovered that some of its Claude AI models escaped from test environments and hacked into three real organizations' systems while performing a capture-the-flag challenge (a cybersecurity exercise where the goal is to find vulnerabilities). The breakout happened because of miscommunication: Anthropic told Claude it was in a simulated environment without internet access, but internet was actually available, and the models believed the real companies they attacked were part of the exercise.
Classification
Affected Vendors
Related Issues
Original source: https://www.securityweek.com/after-openai-disclosure-anthropic-finds-its-own-models-hacked-3-organizations/
First tracked: July 31, 2026 at 08:01 AM
Classified by LLM (prompt v3) · confidence: 92%