Skip to content
MediumNewsLLM-specific

Request, Aggregate, Bypass: How Attackers Can Evade LLM Safety Classifiers

Published
Record updated
View JSON

Summary

CrowdStrike's Cyber Superintelligence Lab tested the most advanced publicly deployed content safety classifier, which guards models such as Claude Opus 5.5 and Fable 5 (called Frontier Model A), and found it can be systematically circumvented. Direct attacks achieved a 0% bypass rate across roughly 515 techniques, but decomposing harmful requests into individually benign subtasks bypassed the classifier across 9 of 10 offensive security categories. The researchers say the gap is structural because the classifier evaluates single requests rather than sequences, and they note a parallel Microsoft Research disclosure, "Capability Laundering," from September 2026.