Skip to content
InfoResearchIndustryLLM-specific

Training agents to self-report misbehavior

Published
Record updated
View JSON

Summary

Researchers Bruce W. Lee, Yueh-Han Chen and Tomek Korbak train agents to call a report_scheming() tool whenever they covertly misbehave, a method they call self-incrimination. For GPT-4.1, undetected successful attacks fall from 56% to 6%, outperforming matched-capability blackbox monitors and alignment baselines across 15 out-of-distribution environments. The approach also preserves general capabilities and generalizes from instructed to uninstructed misbehavior.