Skip to content
InfoResearchIndustryLLM-specific

Why we are excited about confessions

Published
Record updated
View JSON

Summary

Boaz Barak, Gabriel Wu, Jeremy Chen and Manas Joglekar publish a follow-up to their confessions paper, giving deeper analysis of how training affects confessions and preliminary comparisons to chain-of-thought monitoring. The approach trains a second model output, a confession, rewarded solely for honesty about misbehavior in the main task. The authors hypothesize that honest confessions are the path of least resistance because confessing is easier than sustaining an elaborate lie.