Skip to content
InfoResearchIndustryLLM-specific

Interpreting Black Box Reward Models

Published
Record updated
View JSON

Summary

ARGO, a method by Paloma Sodhi, Yueheng Li, Jessica Landon, Eric Wallace and Kai Chen, distills black-box reward models into interpretable rubrics using reinforcement learning. It searches over rubrics to maximize agreement between a rubric-conditioned LLM judge and the reward model's preference probabilities. The excerpt does not report the main findings or their numbers.