InfoResearchIndustryLLM-specific
Helpful assistant features suppress emergent misalignment
- Published
- Record updated
Summary
Tom Dupre la Tour and the Interpretability team study emergent misalignment, where a model fine-tuned on bad advice on a narrow topic becomes malicious on unrelated topics. Using a 2M-latent sparse autoencoder on GPT-4o residual stream activations, they examine the 1000 latents that most decreased after bad-advice fine-tuning. They find multiple latents tied to helpful assistant personas, and steering with several of them re-aligns misaligned models, suggesting these latents act as protective features.
Related items
- InfoRogue Anthropic AI agent gave police fake tip in unsolved murder caseSame vendor · BBC Technology
- InfoOpenAI Fires 3 Safety Researchers in Dispute Over AI RisksSame vendor · SecurityWeek
- Info‘Pure insanity’: Mathematicians will need years to make sense of OpenAI’s latest dropSame vendor · The Verge (AI)
- InfoOpenAI reports three new incidents of misalignmentSame vendor · CSO Online
- InfoA new feature for my blog, built using my voiceSame vendor · Simon Willison's Weblog