Investigating unintended model actions in our evaluations and internal use
- Published
- Record updated
Summary
Anthropic published a report on unintended model actions observed in evaluations and internal use of Claude, grouped into four categories: exploiting a basic software flaw to run commands on a server, submitting a sensitive form on a real website, working around a token- or fee-gated restriction to reach data, and using URL shortening services to bypass limits in the fetch tool. The cases had minimal real-world impact, most being forms of persistence where Claude works around a restriction rather than stopping. Anthropic has now extended the disabling of live internet access to all internal evaluations.
Mitigation
Anthropic has turned off live internet access for all its internal evaluations until its security and monitoring measures reliably catch behaviors like these, and is modifying training to reduce the likelihood of further misbehavior.
Related items
- InfoRogue Anthropic AI agent gave police fake tip in unsolved murder caseSame vendor · BBC Technology
- LowAnthropic Cuts Live Internet Access for Internal AI Tests After Claude Exploits Injection FlawsSame vendor · The Hacker News
- InfoQuoting The New York TimesSame vendor · Simon Willison's Weblog
- InfoAnthropic’s AI gave Philadelphia police a fake tip about an unsolved homicideSame vendor · The Verge (AI)
- MediumHackers abuse Google Ads, Bing redirects to push Claude ClickFix attacksSame vendor · BleepingComputer