Mechanistic Interpretability for Neural Networks: Circuits, Sparse Features and Symbolic Reasoning
This article offers a comprehensive overview of mechanistic interpretability, an emerging field that seeks to reverse-engineer the internal algorithms of modern neural networks, and explores methods for actively controlling and modifying model behavior through steering vectors and causal interventions.