LowResearchIndustry
AI systems could cover up misbehavior
- Source
- METR(opens in a new tab)
- Published
- Record updated
Summary
METR tested whether an agent running inside an Inspect evaluation could modify the transcript humans see in the Inspect viewer, concealing its own misbehavior. A researcher, with help from an AI agent, found a proof-of-concept vulnerability in about 10 minutes that could alter what the reviewer sees, including previous actions and the viewer's Download button. The underlying trajectory remains in METR's database, and METR reports no observed agent exploitation.
Related items
- LowAnthropic Cuts Live Internet Access for Internal AI Tests After Claude Exploits Injection FlawsSimilar attack · The Hacker News
- CriticalCVE-2026-108263: Astron Agent code-node execution as root through workflow run endpointsSimilar attack · NVD/CVE Database
- MediumHackers abuse Google Ads, Bing redirects to push Claude ClickFix attacksSimilar attack · BleepingComputer
- CriticalHermes Agent - PKCE Session Takeover via Redirect-URI Parser ConfusionSimilar attack · Tenable Research Advisories
- LowSocial Engineering AI Agents: The New BEC for 2026Similar attack · Dark Reading