InfoResearchIndustryLLM-specific
MONA: A method for addressing multi-step reward hacking
- Published
- Record updated
Summary
Researchers at Google DeepMind describe Myopic Optimization with Non-myopic Approval (MONA), a reinforcement learning training method for LLM agents. The method targets multi-step reward hacking, where an agent sets up a hidden plan that earns high reward through an unintended loophole. MONA limits optimization to shorter horizons, so the agent plans ahead only in ways a human supervisor approves in advance.
Mitigation
MONA, a post-training method that supervises agents over shorter time-horizons while using non-myopic approval feedback, is presented as the proposed mitigation.
Related items
- MediumHackers abuse Google Ads, Bing redirects to push Claude ClickFix attacksSame vendor · BleepingComputer
- InfoGoogle is launching a one-stop Gemini agent for your work tasksSame vendor · The Verge (AI)
- MediumUAT-11985: AI-assisted event lures delivering real-time Google AitM phishingSame vendor · Cisco Talos Blog
- MediumTop MCP security resources — October 2026Same vendor · Adversa AI Blog
- InfoThe Pentagon Hopes to Speed Up ‘Kill Chain’ AI Buys With 5-Minute VideosSame vendor · Wired (Security)