Recent Test-Time Training (TTT) architectures compress context into fast weights that are updated online and queried as memory. Existing TTT designs keep this memory private to each layer: it recurs only over time, and depth merely indexes L separate memories. We argue that memory ownership need not be tied to depth, a...
Ze-Fan Cai, Qin-Zhe Hu, Ziqiao Ma et al.· 1 citation
Multi-Head Attention Residuals (MHAR) is introduced: the routing query is reshaped into H per-subspace heads, each with its own softmax over the depth history, and a direct probe of the trained queries confirms that learned subspace disagreement is the underlying driver.
Test-Time Training with Next-Token Prediction (TTT-NTP), a drop-in fast-weight adaptation method for pretrained LLMs that instead supervises updates using the model's own next contextual hidden state, while preserving commonsense and knowledge performance.
Xuan Ouyang, Zefan Cai, Junjie Hu· arXiv.org· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.