Skip to content

LoopSpec: Pipelined Self-Speculative Decoding for Looped Transformers

Sep 2026 · 1 citation · 29 references
Computer Science

TL;DR

This work proposes LoopSpec, a training-free self-speculative decoding framework tailored for Looped Transformers that introduces a selective second proposal from deeper recurrent depth while ensuring lossless decoding under both greedy and sampling regimes.

Abstract

Looped Transformers achieve strong performance with compact parameter sizes by repeatedly applying a shared stack of Transformer blocks across recurrent depths. However, they incur higher decoding latency than standard Transformer models of comparable parameter size because shared weights are accessed at every recurrent depth. To improve decoding efficiency, self-speculative decoding is particularly well suited to Looped Transformers, as their intermediate recurrent states can directly provide draft predictions without an auxiliary draft model. We therefore propose LoopSpec, a training-free self-speculative decoding framework tailored for Looped Transformers. LoopSpec extracts draft tokens from early recurrent states and operates in a pipelined manner, overlapping draft generation of future tokens with target verification of the current token. To improve draft accuracy without excessive compute overhead, we introduce a selective second proposal from deeper recurrent depth while ensuring lossless decoding under both greedy and sampling regimes. Furthermore, we derive the optimal proposal depths in closed form and show the prediction matches measurement. Across reasoning and coding benchmarks, LoopSpec achieves up to 6.83$\times$ inference speedup across diverse Looped Transformers.

View source

Similar papers

#machine learning Preprint Oct 2026

Decoding Looped Transformers Better for (Almost) Free

Looped Transformers achieve parameter efficiency by repeatedly executing a shared block across recurrent loops. Each loop yields an intermediate representation decodable for the same next token, yet standard decoding discards earlier states. Because earlier loops embody less computation, recurrence inherently supplies...

Weihao Liu, Huangjie Zheng, Tian-Rong Chen et al. · 0 citations
#machine learning Preprint Sep 2026

WaveFront Decoding: Parallelized Self-Speculative Decoding for Looped Language Models

Looped language models repeatedly apply a weight-shared block to increase effective depth without increasing parameter count, but the resulting T sequential recurrent-block calls per generated token substantially increase decoding latency. To address the issue, we introduce Wavefront Decoding (WFD), a training-free sel...

Hyeongju Ha, Jae-Joon Kim · 0 citations
#machine learning Preprint Sep 2026

Looped Transformers as Optimizers

Looped Transformers provide a parameter-efficient approach to depth scaling by repeatedly applying shared Transformer blocks. Recent reasoning models have likewise highlighted the value of scaling test-time computation through longer computation trajectories. However, the principles for designing effective loop transit...

Yu-Long Huang, Chen Jiang, Zhan-Peng Zhou et al. · 0 citations
#artificial intelligence Preprint Sep 2026

T-LoopFormer: Token-Level Elastic-Depth Looped Transformers for Latent Reasoning with Dynamic Routing

This work proposes dynamic token-choice routing for looped transformers, enabling each token to adaptively determine its own number of loop iterations based on its hidden state, which can improve the token generation accuracy and validate the effectiveness of token-choice router and recursion-wise KV cache.

Ming-Qian Yu, Wen-Peng Zhang, Shao-Bo Cui et al. · 1 citation
#artificial intelligence Preprint Sep 2026

Shallow Queries, Mature Values: Depth-Asynchronous Self-Speculation for Looped Transformers

Looped Transformers reuse a shared block across recurrent depths, making autoregressive decoding expensive because every generated token requires many sequential recurrent passes. Self-speculative decoders reduce this cost by drafting at an early depth and verifying at full depth, but typically bind draft computation t...

Guang-Hao Li, Zi-Han Su, Hao Yu et al. · 0 citations
#machine learning Preprint Sep 2026

FlashLoop: Fast and Memory-Efficient Looped Transformers via Lazy Updates

Looped Transformers have attracted substantial attention as a parameter-efficient approach to increasing computational depth through repeated application of shared Transformer blocks. However, their practical advantages over conventional Transformers remain under debate: each additional loop incurs another Transformer...

Wan-Qi Yang, Shi-Wei Liu · 0 citations

Related blog posts

MIT News · Artificial Intelligence Aug 27, 2026

Looking beyond natural sequences

A new machine-learning framework aims to improve the success rate of computational protein design while moving away from results that reproduce sequences found in nature.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.