Unstable Features, Reproducible Subspaces: Understanding Seed Dependence in Sparse Autoencoders
Some support for the hypothesis that SAE feature instability is caused by the autoencoder tiling a manifold in unique ways. Doesn’t attempt to actually find and describe the manifold, but suggests doing so would be worthwhile.
Structuring Sparsity: Block-Sparse Featurizers Capture Visual Concept Manifolds
Goodfire’s new dictionary learning architecture which captures manifolds and seems like the first one to work on toy models of manifold discovery. Strangely enough, I haven’t seen much discussion of it online, and the release wasn’t publicized. Perhaps an extension to LLMs will be released soon?