{"data":{"id":"eee27e02-4d4a-433a-b235-adbb33af8eca","title":"Request, Aggregate, Bypass: How Attackers Can Evade LLM Safety Classifiers","summary":"CrowdStrike's Cyber Superintelligence Lab tested the most advanced publicly deployed content safety classifier, which guards models such as Claude Opus 5.5 and Fable 5 (called Frontier Model A), and found it can be systematically circumvented. Direct attacks achieved a 0% bypass rate across roughly 515 techniques, but decomposing harmful requests into individually benign subtasks bypassed the classifier across 9 of 10 offensive security categories. The researchers say the gap is structural because the classifier evaluates single requests rather than sequences, and they note a parallel Microsoft Research disclosure, \"Capability Laundering,\" from September 2026.","solution":"N/A -- no mitigation discussed in source.","labels":["security","safety"],"sourceUrl":"https://www.crowdstrike.com/en-us/blog/how-attackers-can-bypass-llm-safety-classifiers/","publishedAt":"2026-10-06T05:00:00.000Z","cveId":null,"cweIds":null,"cvssScore":null,"cvssSeverity":null,"severity":"medium","attackType":["jailbreak","model_evasion"],"issueType":"news","affectedPackages":null,"affectedPackageNames":null,"affectedVendors":["Anthropic"],"affectedVendorsRaw":["Claude Opus 5.5","Fable 5"],"classifierModel":"claude-haiku-5-5","classifierPromptVersion":"v4","summaryPromptVersion":"v2","headline":null,"headlinePromptVersion":null,"cvssVector":null,"attackVector":null,"attackComplexity":null,"privilegesRequired":null,"userInteraction":null,"exploitMaturity":null,"epssScore":null,"epssCheckedAt":null,"kevDateAdded":null,"advisoryAliases":null,"affectedPackagesSource":null,"affectedPackagesCheckedAt":null,"patchAvailable":null,"disclosureDate":"2026-10-06T05:00:00.000Z","capecIds":null,"crossRefCount":0,"attackSophistication":"advanced","impactType":["safety","integrity"],"aiComponentTargeted":"rag","llmSpecific":true,"classifierConfidence":0.85,"researchCategory":null,"atlasIds":null}}