Skip to content
InfoResearchIndustryLLM-specific

Can public chat data predict real-world AI misalignments?

Published
Record updated
View JSON

Summary

Researchers ask whether external groups can evaluate frontier language models by substituting WildChat, a public dataset of about 1 million conversations collected between April 2023 and May 2024, for private production data in their Deployment Simulation technique. They report that WildChat-based predictions of real-world failure rates are surprisingly accurate, typically within roughly 3x error for GPT-5.1, 5.2 and 5.4, despite a 2-3 year data gap. Predictive performance degrades most for more technical and agentic forms of misalignment.