InfoResearchIndustryLLM-specific
Can public chat data predict real-world AI misalignments?
- Published
- Record updated
Summary
Researchers ask whether external groups can evaluate frontier language models by substituting WildChat, a public dataset of about 1 million conversations collected between April 2023 and May 2024, for private production data in their Deployment Simulation technique. They report that WildChat-based predictions of real-world failure rates are surprisingly accurate, typically within roughly 3x error for GPT-5.1, 5.2 and 5.4, despite a 2-3 year data gap. Predictive performance degrades most for more technical and agentic forms of misalignment.
Related items
- InfoRogue Anthropic AI agent gave police fake tip in unsolved murder caseSame vendor · BBC Technology
- InfoOpenAI Fires 3 Safety Researchers in Dispute Over AI RisksSame vendor · SecurityWeek
- Info‘Pure insanity’: Mathematicians will need years to make sense of OpenAI’s latest dropSame vendor · The Verge (AI)
- InfoOpenAI reports three new incidents of misalignmentSame vendor · CSO Online
- InfoA new feature for my blog, built using my voiceSame vendor · Simon Willison's Weblog