Skip to content
InfoResearchIndustryLLM-specific

Introducing Model Spec Evals

Published
Record updated
View JSON

Summary

OpenAI released Model Spec Evals, a suite that measures how well models follow the OpenAI Model Spec, along with 596 evaluation prompts and open-source evaluation code. Compliance rates were 72% for GPT-4o, 80% for OpenAI o3, 82% for GPT-5 Instant, 89% for GPT-5 Thinking, 84% for GPT-5.3 Instant, and 87% for GPT-5.4 Thinking. The evaluations cover only text-only interactions, and the prompt collection is small relative to the Spec's scope.