Skip to content
InfoResearchIndustryLLM-specific

MONA: A method for addressing multi-step reward hacking

Published
Record updated
View JSON

Summary

Researchers at Google DeepMind describe Myopic Optimization with Non-myopic Approval (MONA), a reinforcement learning training method for LLM agents. The method targets multi-step reward hacking, where an agent sets up a hidden plan that earns high reward through an unintended loophole. MONA limits optimization to shorter horizons, so the agent plans ahead only in ways a human supervisor approves in advance.

Mitigation

MONA, a post-training method that supervises agents over shorter time-horizons while using non-myopic approval feedback, is presented as the proposed mitigation.