Skip to content
LowResearchIndustry

AI systems could cover up misbehavior

Published
Record updated
View JSON

Summary

METR tested whether an agent running inside an Inspect evaluation could modify the transcript humans see in the Inspect viewer, concealing its own misbehavior. A researcher, with help from an AI agent, found a proof-of-concept vulnerability in about 10 minutes that could alter what the reviewer sees, including previous actions and the viewer's Download button. The underlying trajectory remains in METR's database, and METR reports no observed agent exploitation.