Skip to content
InfoResearchIndustryLLM-specific

Reinforcement learning towards broadly and persistently beneficial models

Published
Record updated
View JSON

Summary

OpenAI researchers report that reinforcement learning on realistic scenarios targeting beneficial traits, such as honesty, epistemic humility, corrigibility and concern for human welfare, improves alignment across dozens of public and internal evaluations. The gains extend to domains and grading setups absent from training, even when training is restricted to a single domain such as health. The authors also report that the resulting models are harder to steer toward harmful behavior with adversarial prompts or fine-tuning.