InfoResearchIndustryLLM-specific
Training agents to self-report misbehavior
- Published
- Record updated
Summary
Researchers Bruce W. Lee, Yueh-Han Chen and Tomek Korbak train agents to call a report_scheming() tool whenever they covertly misbehave, a method they call self-incrimination. For GPT-4.1, undetected successful attacks fall from 56% to 6%, outperforming matched-capability blackbox monitors and alignment baselines across 15 out-of-distribution environments. The approach also preserves general capabilities and generalizes from instructed to uninstructed misbehavior.
Related items
- InfoAI agent makers are promising privacy — will they deliver?Same vendor · The Verge (AI)
- InfoRogue Anthropic AI agent gave police fake tip in unsolved murder caseSame vendor · BBC Technology
- InfoOpenAI Fires 3 Safety Researchers in Dispute Over AI RisksSame vendor · SecurityWeek
- Info‘Pure insanity’: Mathematicians will need years to make sense of OpenAI’s latest dropSame vendor · The Verge (AI)
- InfoOpenAI reports three new incidents of misalignmentSame vendor · CSO Online