InfoResearchIndustryLLM-specific
Investigating the consequences of accidentally grading CoT during RL
- Published
- Record updated
Summary
OpenAI's researchers report that an automated system found limited accidental Chain-of-Thought (CoT) grading during RL training of some released models, including GPT-5.4 Thinking, GPT-5.1 Instant through GPT-5.4 Instant, GPT-5.3 mini, and GPT-5.4 mini, despite the company's policy against it. Their analysis found no clear reduction in CoT monitorability, though they cannot rule out harder-to-measure effects.
Mitigation
Fixed the affected reward pathways, expanding the automated detection system, and strengthened internal guidance and processes.
Related items
- InfoRogue Anthropic AI agent gave police fake tip in unsolved murder caseSame vendor · BBC Technology
- InfoOpenAI Fires 3 Safety Researchers in Dispute Over AI RisksSame vendor · SecurityWeek
- Info‘Pure insanity’: Mathematicians will need years to make sense of OpenAI’s latest dropSame vendor · The Verge (AI)
- InfoOpenAI reports three new incidents of misalignmentSame vendor · CSO Online
- InfoA new feature for my blog, built using my voiceSame vendor · Simon Willison's Weblog