Skip to content
InfoResearchIndustryLLM-specific

Consistency Training Could Help Limit Sycophancy and Jailbreaks

Published
Record updated
View JSON

Summary

Authors Alex Irpan, Alex Turner, Mark Kurzeja, David Elson and Rohin Shah propose consistency training, a self-supervised method that teaches a model to ignore irrelevant cues such as user biases or jailbreak wrappers. They introduce Activation Consistency Training (ACT), which optimizes internal activations, and compare it with Bias-augmented Consistency Training (BCT) and baselines on Gemma 2, Gemma 3 and Gemini 2.5 Flash, reporting that ACT and BCT beat the baselines and improve robustness. The approach avoids the static datasets that SFT relies on, which can go stale in guidelines or capability.