Skip to content

RunningTensor: Generalizing Linear Attention to Higher-Order Recurrent States

Sep 2026 · 1 citation · 44 references
Computer Science

TL;DR

The RunningTensor is introduced, which generalizes this memory to an order-o tensor, updated by a rank-1 outer product and read by contracting against $o-1$ vector queries, and improves performance on language-understanding and non-synthetic retrieval tasks, suggesting that higher-order recurrent state can provide useful additional memory capacity beyond matrix-valued state.

Abstract

Linear attention and state-space models provide linear-time sequence modeling, but their recurrent memory remains a second-order tensor (a matrix), limiting the order of interactions that can be represented in the state. We introduce the RunningTensor, which generalizes this memory to an order-$o$ tensor, updated by a rank-1 outer product and read by contracting against $o-1$ vector queries. Order $2$ recovers linear attention; we study order $3$ as a proof of concept, retaining both recurrent and parallel forms while remaining linear in sequence length $T$ and improving working memory capacity from $\mathcal{O}(W^2)$ to $\mathcal{O}(W^o)$. On synthetic multi-query associative recall, RunningTensor outperforms linear-attention and SSM baselines. After pretraining, it also improves performance on language-understanding and non-synthetic retrieval tasks, suggesting that higher-order recurrent state can provide useful additional memory capacity beyond matrix-valued state.

View source

Similar papers

#machine learning Preprint Sep 2026

Triadic Linear Attention: Three-Dimensional Recurrent States for Long-Context Sequence Modeling

Recurrent neural networks (RNNs) compress the historical context into a memory state of fixed size, thus allowing for constant-time inference. The memory state size is a crucial factor in their performance, as exemplified by the strong performance and resurgence of linear attention, which extends the vector-valued hidd...

Oliver Sieberling, Bharat Runwal, David Jin et al. · 0 citations
#machine learning Preprint Aug 2026

Fast Weight Attention for Continual Learning

This framework separates temporal alignment, plasticity, forgetting, and bounded rehearsal in recurrent sequence models, together with numerically stable positive-decay renormalization, to remain competitive in language modeling and improve length extrapolation on variable-digit addition.

Yi-Fan Zhang, Steve Ta, Jasper Zhang et al. · 0 citations
#artificial intelligence Preprint Sep 2026

SMat-Attention: Structured Long-Context Sequence Modeling

Long-context sequence models face a fundamental tradeoff: softmax attention uses flexible token-level interactions at quadratic cost, whereas linear attention obtains linear-time training and constant-time decoding by compressing history into a fixed-size state. In this work, we ask whether we can connect these regimes...

E. Anand, Abdullah Ateyeh, Archer Wang et al. · 2 citations
#machine learning Preprint Sep 2026

CyFA: Linear Sequence Modeling with Relative-Time-Partitioned Memory

Linear RNNs offer linear-time sequence processing and constant-memory decoding, but their fixed-size recurrent states must accommodate all past key--value associations. Existing forgetting mechanisms and Delta Rule updates reduce interference by selectively clearing or correcting the state, yet earlier associations can...

Yi-Xiao Chen, Shuo-Jin Yang, Shi-Min Hu · 0 citations
#artificial intelligence Preprint Sep 2026

Switching Linear Attention

Designing expressive sequence layers with efficient inference remains a central challenge in modern machine learning. Standard softmax attention achieves excellent sequence modeling performance through rich nonlinear token interactions, but it requires a key-value cache that grows linearly with sequence length, limitin...

Hyun Dong Lee, X. Gonzalez, Nicolas Zucchet et al. · 0 citations
#machine learning Preprint Sep 2026

High-Dimensional Learning Dynamics of Attention-Indexed Models

This work studies attention-indexed models, a broad framework that can represent multi-layer and multi-head attention architectures, and reveals that attention parameterization itself can act as an architectural implicit bias.

Yi-Zhou Xu, M. Sagitova, L. Zdeborová et al. · 0 citations

Related blog posts

MIT News · Artificial Intelligence Sep 29, 2026

Who we become when we talk to machines

Professor Sherry Turkle’s new book, “Artificial Intimacy,” offers a withering critique of chatbots and the antisocial dynamics she believes they encourage.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.