InfoResearchPreprintLLM-specific
PatchBench: Measuring Collateral Damage in Activation Patching
- Published
- Record updated
Summary
PatchBench introduces a benchmark of 400 curated, model-specific jailbreak failures drawn from 27,870 prompts across 37 public datasets, filtered with WildGuard, pairwise Elo ranking and manual verification. Its companion protocol, PatchBench-Local, tests whether a patch is behaviourally precise by checking harmful-neighbour correction and benign-neighbour preservation. Evaluating four activation steering methods, the authors found global capability could stay nearly unchanged while local benign regressions were severe, showing aggregate metrics miss collateral damage.
Related items
- CriticalHermes Agent - PKCE Session Takeover via Redirect-URI Parser ConfusionSimilar attack · Tenable Research Advisories
- LowLost in the comments: Social context as a single‐pass jailbreak and defense on agentic platformsSimilar attack · OpenAlex (peer-reviewed AI security)
- MediumGHSA-hmq2-7hp6-7crh: Banks: User-controlled prompt input can be parsed as privileged chat messagesSimilar attack · GitHub Advisory Database
- HighGHSA-6wjp-v33h-5cvq: PraisonAI: AgentOS defaults to network-exposed no-auth mode, allowing unauthenticated agent invocation and instruction disclosureSimilar attack · GitHub Advisory Database
- HighCVE-2026-101998: Docker Sandboxes fail open when masking credentials in proxy responsesSimilar attack · NVD/CVE Database