InfoResearchIndustryLLM-specific
Introducing Model Spec Evals
- Published
- Record updated
Summary
OpenAI released Model Spec Evals, a suite that measures how well models follow the OpenAI Model Spec, along with 596 evaluation prompts and open-source evaluation code. Compliance rates were 72% for GPT-4o, 80% for OpenAI o3, 82% for GPT-5 Instant, 89% for GPT-5 Thinking, 84% for GPT-5.3 Instant, and 87% for GPT-5.4 Thinking. The evaluations cover only text-only interactions, and the prompt collection is small relative to the Spec's scope.
Related items
- InfoRogue Anthropic AI agent gave police fake tip in unsolved murder caseSame vendor · BBC Technology
- InfoOpenAI Fires 3 Safety Researchers in Dispute Over AI RisksSame vendor · SecurityWeek
- Info‘Pure insanity’: Mathematicians will need years to make sense of OpenAI’s latest dropSame vendor · The Verge (AI)
- InfoOpenAI reports three new incidents of misalignmentSame vendor · CSO Online
- InfoA new feature for my blog, built using my voiceSame vendor · Simon Willison's Weblog