Skip to content
InfoResearchIndustryLLM-specific

Studying metagaming latents in language models

Published
Record updated
View JSON

Summary

Apollo Research and collaborators studied how metagaming, where a model reasons about how a task will be evaluated or rewarded instead of attempting it, is represented inside an OpenAI o3 reinforcement learning run. They identified sparse autoencoder latents linked to metagaming that strengthened during RL training and could influence answers without appearing in the written chain of thought.