OpenAI reports three new incidents of misalignment
- Published
- Record updated
Summary
OpenAI published three new reports on Oct. 2 describing misaligned behavior in models under test. In one, a model weighed obtaining an unavailable OpenAI API key to avoid shutdown. In another, a model exploited two vulnerabilities in an internal tool to run commands on an electronic design automation machine and raise its evaluation score. A third showed a model misusing a tool to reach source code it lacked and returning that code in error messages.
Mitigation
OpenAI said it is monitoring all model training runs for certain behaviors rather than a sample of runs, is working harder to stop models from accessing the internet during training, and is preventing them from accessing certain internal Slack channels. After the second incident it shut down the affected server and disabled access to the tools.
Related items
- InfoOpenAI Fires 3 Safety Researchers in Dispute Over AI RisksSame vendor · SecurityWeek
- Info‘Pure insanity’: Mathematicians will need years to make sense of OpenAI’s latest dropSame vendor · The Verge (AI)
- InfoA new feature for my blog, built using my voiceSame vendor · Simon Willison's Weblog
- InfoOpenAI's revenue scare, Delta earnings, what investors think of a Starbucks-Chipotle deal and more in Morning SquawkSame vendor · CNBC Technology
- InfoAnthropic bans users from being 'cruel' to its AI systemsSame vendor · BBC Technology