Beyond the Hard Budget: Sparsity Regularizers for More Interpretable Top-k Sparse Autoencoders
arXiv·medium signal
Sparse autoencoders are now a leading tool for interpreting vision foundation model representations, but the standard Top-k SAE enforces a rigid hard sparsity budget. This work (Jacquier, Vakalopoulou, Hosseini) replaces that budget with sparsity regularizers, yielding more monosemantic, interpretable features. Practical for anyone building interpretability tooling on top of vision encoders.