Skip to content
InfoResearchPreprintLLM-specific

PatchBench: Measuring Collateral Damage in Activation Patching

Published
Record updated
View JSON

Summary

PatchBench introduces a benchmark of 400 curated, model-specific jailbreak failures drawn from 27,870 prompts across 37 public datasets, filtered with WildGuard, pairwise Elo ranking and manual verification. Its companion protocol, PatchBench-Local, tests whether a patch is behaviourally precise by checking harmful-neighbour correction and benign-neighbour preservation. Evaluating four activation steering methods, the authors found global capability could stay nearly unchanged while local benign regressions were severe, showing aggregate metrics miss collateral damage.