Local Sparsity Enables Unsupervised LLM Safety Detection
This work proposes a framework for locally masked SAE-based anomaly detection, supported by theoretical justifications, and demonstrates their ability to capture meaningful safety information while using only 1-2% of SAE neurons for computation.