Why Knowing Both Hops Is Not Enough: Understanding Two-Hop Generalization in Language Models
A recurrent-style training strategy is proposed, which enables transformers to reuse their reasoning circuitry across input forms and substantially improves generalization on out-of-distribution two-hop queries.
Zili Zhang, Yilin Wang, Heng Wang et al.
· 0 citations