Skip to content
InfoResearchIndustryLLM-specific

Investigating the consequences of accidentally grading CoT during RL

Published
Record updated
View JSON

Summary

OpenAI's researchers report that an automated system found limited accidental Chain-of-Thought (CoT) grading during RL training of some released models, including GPT-5.4 Thinking, GPT-5.1 Instant through GPT-5.4 Instant, GPT-5.3 mini, and GPT-5.4 mini, despite the company's policy against it. Their analysis found no clear reduction in CoT monitorability, though they cannot rule out harder-to-measure effects.

Mitigation

Fixed the affected reward pathways, expanding the automated detection system, and strengthened internal guidance and processes.