Skip to content
InfoResearchIndustryLLM-specific

Discovering unknown AI misalignments in real-world usage

Published
Record updated
View JSON

Summary

OpenAI describes an approach that uses reasoning models as AI judges to find misaligned behavior in real-world ChatGPT conversations by detecting sentiment deterioration in users. The judges analyze historic production conversations where users allowed their data to be used, and clustering identifies common themes. Conversations with sentiment deterioration were roughly twice as likely to contain OpenAI Model Spec violations.