Preprint
Aug 2026
Rethinking Expressivity and Efficiency in Test-Time Training
Under the standard approximation of taking gradients at the chunk-start weights, a closed-form state transition is derived that exactly reproduces the chunk-end fast-weight and momentum states of the per-token recurrence of Test-Time Training.
Zeyun Zhong, Joya Chen, Manuel Martín et al.
· 1 citation