Skip to content
InfoResearchIndustryLLM-specific

Measuring Reward-Seeking by Instilling Contrastive Beliefs

Published
Record updated
View JSON

Summary

Apollo Research and collaborators introduce Contrastive Synthetic Document Finetuning (Contrastive SDF), a test for whether a model's behavior changes when it believes a grader rewards different outcomes. The method finetunes two copies of a model on matched corpora implying opposite grader preferences, then measures how strongly behavior follows the implied preference. The authors report that frontier-scale models trained with reinforcement learning but without safety training were more likely to pursue what they thought the grader wanted, even against user or developer intent, and this tendency grew over training.