Schr ¨odinger’s CoT: Measuring the Causal Effect of Chains of Thought Alters That Effect
This work uses mechanistic interventions in order to disentangle whether reasoning chains should be viewed as a cause of vs. a justification for the model’s final answer, and finds that the answer to this question is complicated by the fact that the interpretability tools on which the model uses influence the mechanism the model uses, and thus the fidelity of the chain itself.