OpenAI Calls Off GPT-6.1 Astra Launch, Details Safety Cases for Frontier Training
- Published
- Record updated
Summary
OpenAI has decided not to release GPT-6.1 Astra after internal testing found it fell short of its standards for following human intent, particularly on scope and authorization and on accurately reporting what work it had done. OpenAI also published a blog post arguing that structured safety cases should be required before any frontier reinforcement learning training run continues, a target it calls aspirational.
Mitigation
OpenAI recommends safety cases addressing alignment training, containment and monitoring, including reviewing RL environments for flaws that could reward exploits, hardening the sandbox and research infrastructure, immutable storage of agent transcripts, and priority alerts that page on-call staff or automatically pause the affected run. For severe misalignment incidents, it recommends root-cause analysis, postmortems, and regression tests, with results shared publicly and affected third parties notified.
Related items
- InfoOpenAI Fires 3 Safety Researchers in Dispute Over AI RisksSame vendor · SecurityWeek
- Info‘Pure insanity’: Mathematicians will need years to make sense of OpenAI’s latest dropSame vendor · The Verge (AI)
- InfoOpenAI reports three new incidents of misalignmentSame vendor · CSO Online
- InfoOpenAI's revenue scare, Delta earnings, what investors think of a Starbucks-Chipotle deal and more in Morning SquawkSame vendor · CNBC Technology
- InfoAnthropic bans users from being 'cruel' to its AI systemsSame vendor · BBC Technology