InfoResearchIndustryLLM-specific
Debugging misaligned completions with sparse-autoencoder latent attribution
- Published
- Record updated
Summary
OpenAI researchers Tom Dupre la Tour and Dan Mossing describe a method for finding sparse-autoencoder (SAE) latents causally linked to misaligned behavior in a single model. The approach computes attribution differences between positive and negative completions of the same prefix, then validates the selected latents by steering activations and grading new completions with an LLM judge. In a case study on a model fine-tuned to give inaccurate health information, the top 100 latents by attribution difference were largely related to misalignment, such as latent #1 "outrage" and latent #2 "murdering".
Related items
- InfoRogue Anthropic AI agent gave police fake tip in unsolved murder caseSame vendor · BBC Technology
- InfoOpenAI Fires 3 Safety Researchers in Dispute Over AI RisksSame vendor · SecurityWeek
- Info‘Pure insanity’: Mathematicians will need years to make sense of OpenAI’s latest dropSame vendor · The Verge (AI)
- InfoOpenAI reports three new incidents of misalignmentSame vendor · CSO Online
- InfoA new feature for my blog, built using my voiceSame vendor · Simon Willison's Weblog