Jul 2026
Looped Latent Attention: Cross-Loop KV Compression for Looped Transformers
This work introduces Looped Latent Attention (\lla{}), a post-training cache codec that stores compact K and V latents and reconstructs loop-specific K/V vectors only when attention reads them, and shows that the recurrent cache is low-rank but not safely collapsible to a single state.
James O'Neill, Fergal Reid
· arXiv.org · 0 citations