Skip to content
InfoResearchIndustryLLM-specific

Negative Results for Sparse Autoencoders On Downstream Tasks and Deprioritising SAE Research…

Published
Record updated
View JSON

Summary

Google DeepMind's mechanistic interpretability team tested whether sparse autoencoders (SAEs) help with out-of-distribution detection of harmful intent in user prompts. SAEs underperformed linear probes, which the team found cheap and strong. As a result, the team is deprioritising fundamental SAE research while keeping SAEs as one tool.