InfoResearchPreprintLLM-specific
BRANCH: Bypassing Multi-Scanner AI Guardrails
- Published
- Record updated
Summary
Researchers present BRANCH, a bypass method for guardrail systems built from multiple scanners that detect malicious instructions aimed at LLMs. The method uses a branching tree search that applies adversarial perturbations against individual scanners and selects techniques by overall improvement across all scanners. The authors report a 100% attack success rate across 6 guardrail systems in 120 scenarios, using 72% fewer queries and 4.5x less wallclock time than established techniques, and note that bypasses transfer to 29 unseen guardrails, including 8 commercial black-box ones.
Related items
- CriticalHermes Agent - PKCE Session Takeover via Redirect-URI Parser ConfusionSimilar attack · Tenable Research Advisories
- LowSocial Engineering AI Agents: The New BEC for 2026Similar attack · Dark Reading
- LowLost in the comments: Social context as a single‐pass jailbreak and defense on agentic platformsSimilar attack · OpenAlex (peer-reviewed AI security)
- MediumGHSA-hmq2-7hp6-7crh: Banks: User-controlled prompt input can be parsed as privileged chat messagesSimilar attack · GitHub Advisory Database
- HighGHSA-6wjp-v33h-5cvq: PraisonAI: AgentOS defaults to network-exposed no-auth mode, allowing unauthenticated agent invocation and instruction disclosureSimilar attack · GitHub Advisory Database