InfoResearchIndustryLLM-specific
Discovering unknown AI misalignments in real-world usage
- Published
- Record updated
Summary
OpenAI describes an approach that uses reasoning models as AI judges to find misaligned behavior in real-world ChatGPT conversations by detecting sentiment deterioration in users. The judges analyze historic production conversations where users allowed their data to be used, and clustering identifies common themes. Conversations with sentiment deterioration were roughly twice as likely to contain OpenAI Model Spec violations.
Related items
- InfoRogue Anthropic AI agent gave police fake tip in unsolved murder caseSame vendor · BBC Technology
- InfoOpenAI Fires 3 Safety Researchers in Dispute Over AI RisksSame vendor · SecurityWeek
- Info‘Pure insanity’: Mathematicians will need years to make sense of OpenAI’s latest dropSame vendor · The Verge (AI)
- InfoOpenAI reports three new incidents of misalignmentSame vendor · CSO Online
- InfoA new feature for my blog, built using my voiceSame vendor · Simon Willison's Weblog