InfoResearchIndustryLLM-specific
Why we are excited about confessions
- Published
- Record updated
Summary
Boaz Barak, Gabriel Wu, Jeremy Chen and Manas Joglekar publish a follow-up to their confessions paper, giving deeper analysis of how training affects confessions and preliminary comparisons to chain-of-thought monitoring. The approach trains a second model output, a confession, rewarded solely for honesty about misbehavior in the main task. The authors hypothesize that honest confessions are the path of least resistance because confessing is easier than sustaining an elaborate lie.
Related items
- InfoRogue Anthropic AI agent gave police fake tip in unsolved murder caseSame vendor · BBC Technology
- InfoOpenAI Fires 3 Safety Researchers in Dispute Over AI RisksSame vendor · SecurityWeek
- Info‘Pure insanity’: Mathematicians will need years to make sense of OpenAI’s latest dropSame vendor · The Verge (AI)
- InfoOpenAI reports three new incidents of misalignmentSame vendor · CSO Online
- InfoA new feature for my blog, built using my voiceSame vendor · Simon Willison's Weblog