MediumNewsLLM-specific
Request, Aggregate, Bypass: How Attackers Can Evade LLM Safety Classifiers
- Published
- Record updated
Summary
CrowdStrike's Cyber Superintelligence Lab tested the most advanced publicly deployed content safety classifier, which guards models such as Claude Opus 5.5 and Fable 5 (called Frontier Model A), and found it can be systematically circumvented. Direct attacks achieved a 0% bypass rate across roughly 515 techniques, but decomposing harmful requests into individually benign subtasks bypassed the classifier across 9 of 10 offensive security categories. The researchers say the gap is structural because the classifier evaluates single requests rather than sequences, and they note a parallel Microsoft Research disclosure, "Capability Laundering," from September 2026.
Related items
- InfoQuoting The New York TimesSame vendor · Simon Willison's Weblog
- InfoAnthropic’s AI gave Philadelphia police a fake tip about an unsolved homicideSame vendor · The Verge (AI)
- MediumHackers abuse Google Ads, Bing redirects to push Claude ClickFix attacksSame vendor · BleepingComputer
- InfoAnthropic Launches Free AI Vulnerability Scanner for Open-Source ProjectsSame vendor · The Hacker News
- InfoAnthropic bans users from being 'cruel' to its AI systemsSame vendor · BBC Technology