State tracking requires an update at every input, but the depth a Transformer applies to each token is fixed regardless of sequence length. We introduce the Recurrent Looped Transformer (RLT), which splits its layers between a parallel causal encoder and a recurrent decoder. At each token, the decoder merges the encode...
This work pretrain decoder-only models with approximately GPT-2 small and medium sizes on the FineWeb-Edu dataset and introduces Deep Delta Learning (DDL), which applies the delta rule over network depth.
Yi-Fan Zhang, Yi-Feng Liu, Mengdi Wang et al.· arXiv.org· 8 citations· ⚡2
The results show that stable recurrent depth requires residual scaling rules that account for parameter visits, not only nominal layer count, and DeepLoop is neutral when no physical block is revisited and improves validation loss and downstream accuracy once recurrent depth is activated.