Shallow Queries, Mature Values: Depth-Asynchronous Self-Speculation for Looped Transformers
Looped Transformers reuse a shared block across recurrent depths, making autoregressive decoding expensive because every generated token requires many sequential recurrent passes. Self-speculative decoders reduce this cost by drafting at an early depth and verifying at full depth, but typically bind draft computation t...