#machine learning
Jul 2026
Are Single-Token Sparse Autoencoder Features Causally Necessary? Layer-Depth and SAE-Family Effects
This work analyzesparse autoencoder features across six models and three SAE families and zero-ablate at full layer depth, finding cross-family claims are sensitive to training methodology, not just activation function or scale.
Seonglae Cho, Zekun Wu, Kleyton Da Costa et al.
· arXiv.org · 1 citation