InfoResearchIndustryLLM-specific
Consistency Training Could Help Limit Sycophancy and Jailbreaks
- Published
- Record updated
Summary
Authors Alex Irpan, Alex Turner, Mark Kurzeja, David Elson and Rohin Shah propose consistency training, a self-supervised method that teaches a model to ignore irrelevant cues such as user biases or jailbreak wrappers. They introduce Activation Consistency Training (ACT), which optimizes internal activations, and compare it with Bias-augmented Consistency Training (BCT) and baselines on Gemma 2, Gemma 3 and Gemini 2.5 Flash, reporting that ACT and BCT beat the baselines and improve robustness. The approach avoids the static datasets that SFT relies on, which can go stale in guidelines or capability.
Topics
Related items
- MediumHackers abuse Google Ads, Bing redirects to push Claude ClickFix attacksSame vendor · BleepingComputer
- InfoGoogle is launching a one-stop Gemini agent for your work tasksSame vendor · The Verge (AI)
- MediumUAT-11985: AI-assisted event lures delivering real-time Google AitM phishingSame vendor · Cisco Talos Blog
- MediumTop MCP security resources — October 2026Same vendor · Adversa AI Blog
- InfoThe Pentagon Hopes to Speed Up ‘Kill Chain’ AI Buys With 5-Minute VideosSame vendor · Wired (Security)