InfoResearchIndustryLLM-specific
Open Sourcing Monitorability Evaluations
- Published
- Record updated
Summary
OpenAI released datasets and reference code from its Monitoring Monitorability paper, covering most of its chain-of-thought monitorability evaluation suite, code for computing the g-mean 2 metric, and a cross-fit filtering strategy for noise-dominated intervention instances. Several evals were omitted because they rely on private or restricted data. The scaffold included in the release is an illustration only, and OpenAI will not support it.
Related items
- InfoRogue Anthropic AI agent gave police fake tip in unsolved murder caseSame vendor · BBC Technology
- InfoOpenAI Fires 3 Safety Researchers in Dispute Over AI RisksSame vendor · SecurityWeek
- Info‘Pure insanity’: Mathematicians will need years to make sense of OpenAI’s latest dropSame vendor · The Verge (AI)
- InfoOpenAI reports three new incidents of misalignmentSame vendor · CSO Online
- InfoA new feature for my blog, built using my voiceSame vendor · Simon Willison's Weblog