A Unifying Perspective on Language Model Representations: From Filler-Role Structure to Mechanistic Interpretability
This work proposes using Tensor Product Representations (TPRs) as a unifying hypothesis, and shows that TPRs can unify several prior interpretability methods: additive analogies, linear probing, sparse autoencoders, and activation patching.
Enshang Zhang, R. Thomas McCoy
· 0 citations