Skip to content
InfoNewsLLM-specific

Quoting Anthropic Frontier Red Team

Published
Record updated
View JSON

Summary

Anthropic's Frontier Red Team evaluated several models on 100 randomly selected tasks from an internal Binary Exploitation benchmark, dated 29 September 2026. GLM-5.3 achieved full control flow hijacks in 4% of trials and Claude Mythos Preview in 6%, while earlier models such as Claude Opus 4.6 and GLM-5.2 succeeded in none.