Sparse Autoencoders: Real Features or Just a Fancy Illusion?
New research reveals that up to 77% of features in degraded Sparse Autoencoders (SAEs) are functionally inert, and even in well-trained models 9% of matched features are dead, challenging the interpre…