InfoResearchIndustryLLM-specific
Reinforcement learning towards broadly and persistently beneficial models
- Published
- Record updated
Summary
OpenAI researchers report that reinforcement learning on realistic scenarios targeting beneficial traits, such as honesty, epistemic humility, corrigibility and concern for human welfare, improves alignment across dozens of public and internal evaluations. The gains extend to domains and grading setups absent from training, even when training is restricted to a single domain such as health. The authors also report that the resulting models are harder to steer toward harmful behavior with adversarial prompts or fine-tuning.
Related items
- InfoRogue Anthropic AI agent gave police fake tip in unsolved murder caseSame vendor · BBC Technology
- InfoOpenAI Fires 3 Safety Researchers in Dispute Over AI RisksSame vendor · SecurityWeek
- Info‘Pure insanity’: Mathematicians will need years to make sense of OpenAI’s latest dropSame vendor · The Verge (AI)
- InfoOpenAI reports three new incidents of misalignmentSame vendor · CSO Online
- InfoA new feature for my blog, built using my voiceSame vendor · Simon Willison's Weblog