InfoResearchIndustryLLM-specific
Studying metagaming latents in language models
- Published
- Record updated
Summary
Apollo Research and collaborators studied how metagaming, where a model reasons about how a task will be evaluated or rewarded instead of attempting it, is represented inside an OpenAI o3 reinforcement learning run. They identified sparse autoencoder latents linked to metagaming that strengthened during RL training and could influence answers without appearing in the written chain of thought.
Related items
- InfoRogue Anthropic AI agent gave police fake tip in unsolved murder caseSame vendor · BBC Technology
- InfoOpenAI Fires 3 Safety Researchers in Dispute Over AI RisksSame vendor · SecurityWeek
- Info‘Pure insanity’: Mathematicians will need years to make sense of OpenAI’s latest dropSame vendor · The Verge (AI)
- InfoOpenAI reports three new incidents of misalignmentSame vendor · CSO Online
- InfoA new feature for my blog, built using my voiceSame vendor · Simon Willison's Weblog