Thinking vs. NoThinking: Towards Interpreting Reasoning Mechanisms of Large Language Models via Sparse Autoencoders
This work applies Top-K Sparse Autoencoders to the intermediate representations of DeepSeek-R1-Distill-Qwen-7B and examines the model's divergent behaviors across math-solving tasks of three distinct difficulty levels, identifying a clear distinction in how the model functions under two reasoning modes.