Skip to content
InfoResearchIndustryLLM-specific

Debugging misaligned completions with sparse-autoencoder latent attribution

Published
Record updated
View JSON

Summary

OpenAI researchers Tom Dupre la Tour and Dan Mossing describe a method for finding sparse-autoencoder (SAE) latents causally linked to misaligned behavior in a single model. The approach computes attribution differences between positive and negative completions of the same prefix, then validates the selected latents by steering activations and grading new completions with an LLM judge. In a case study on a model fine-tuned to give inaccurate health information, the top 100 latents by attribution difference were largely related to misalignment, such as latent #1 "outrage" and latent #2 "murdering".