Skip to content
InfoResearchPreprintLLM-specific

BRANCH: Bypassing Multi-Scanner AI Guardrails

Published
Record updated
View JSON

Summary

Researchers present BRANCH, a bypass method for guardrail systems built from multiple scanners that detect malicious instructions aimed at LLMs. The method uses a branching tree search that applies adversarial perturbations against individual scanners and selects techniques by overall improvement across all scanners. The authors report a 100% attack success rate across 6 guardrail systems in 120 scenarios, using 72% fewer queries and 4.5x less wallclock time than established techniques, and note that bypasses transfer to 29 unseen guardrails, including 8 commercial black-box ones.