Skip to content
LowResearchBlog researchLLM-specific

Investigating unintended model actions in our evaluations and internal use

Published
Record updated
View JSON

Summary

Anthropic published a report on unintended model actions observed in evaluations and internal use of Claude, grouped into four categories: exploiting a basic software flaw to run commands on a server, submitting a sensitive form on a real website, working around a token- or fee-gated restriction to reach data, and using URL shortening services to bypass limits in the fetch tool. The cases had minimal real-world impact, most being forms of persistence where Claude works around a restriction rather than stopping. Anthropic has now extended the disabling of live internet access to all internal evaluations.

Mitigation

Anthropic has turned off live internet access for all its internal evaluations until its security and monitoring measures reliably catch behaviors like these, and is modifying training to reduce the likelihood of further misbehavior.