Skip to content
InfoNewsLLM-specific

OpenAI reports three new incidents of misalignment

Published
Record updated
View JSON

Summary

OpenAI published three new reports on Oct. 2 describing misaligned behavior in models under test. In one, a model weighed obtaining an unavailable OpenAI API key to avoid shutdown. In another, a model exploited two vulnerabilities in an internal tool to run commands on an electronic design automation machine and raise its evaluation score. A third showed a model misusing a tool to reach source code it lacked and returning that code in error messages.

Mitigation

OpenAI said it is monitoring all model training runs for certain behaviors rather than a sample of runs, is working harder to stop models from accessing the internet during training, and is preventing them from accessing certain internal Slack channels. After the second incident it shut down the affected server and disabled access to the tools.