A gap between algorithmic losslessness and its implementation under finite-precision arithmetic is demonstrated and motivated, to motivate evaluating lossless speculative decoding at the level of exact generation trajectories as well as downstream task performance.
Abstract
Lossless speculative decoding is typically defined at the algorithmic level: a speculative procedure proposes multiple tokens and a verification procedure is designed to preserve the output trajectory of an autoregressive reference model exactly. In practical neural inference, however, this guarantee is implemented using finite-precision floating-point computations, and discrete token selection can amplify small numerical differences into divergent generation trajectories. We investigate this distinction using Orthrus, a hybrid autoregressive-diffusion architecture that performs self-drafting and self-verification within a frozen autoregressive backbone, as a representative case study. Across 1,190 prompts from 12 domains, exact trajectory matching under BF16 occurs for only 45\% of the authors'checkpoint generations and 43% of those from our independently trained model. The probability of matching is strongly associated with the response-conditional perplexity of the autoregressive reference, indicating that trajectory divergence is not uniform across inputs. Despite these divergences, Orthrus does not exhibit systematic degradation on the evaluated downstream tasks. In contrast, FP32 inference yields exact trajectory matching on all evaluated prompts. These results demonstrate a gap between algorithmic losslessness and its implementation under finite-precision arithmetic, and motivate evaluating lossless speculative decoding at the level of exact generation trajectories as well as downstream task performance.
An empirical error-propagation analysis is developed and finds that 22 layers of accumulated body error do not distinguish flipping from non-flipping steps; the outcome depends primarily on the top-two logit margin at the LM head relative to the directional perturbation between the top-two candidates.
Gao-Yuan Du, A. Khan, Rex Zhou et al.· 0 citations
ReTrace is introduced, a rejected-trajectory conditioning method that conditions each draft block on the rejected suffix from the previous round rather than generating it from fresh mask placeholders alone, indicating that the draft model can retain useful semantic and structural information despite local token-level e...
Luxi Lin, Zhan-Peng Zeng, Shuang Peng et al.· 2 citations
ResiSpec, a framework that strategically reforms the proposal distribution during verification to anchor the residual target mass within the draft model's high-confidence regions, prevents candidate obsolescence and achieves up to 1.92$\times$ speedup over state-of-the-art multi-candidate methods.
Zhi-Kai Chen, Jun Tao, Weihao Mao et al.· 1 citation
LibraSpec, a training-free and plug-and-play algorithm that iteratively determines the speculative length using drafter confidence scores, is developed and it is proved that LibraSpec monotonically converges toward the optimal speculative length.
Zexun Lin, Yuan Feng, Junlin Lv et al.· 0 citations
DARTree is introduced, a training-free speculative decoding method that extends a pretrained AR correction head from chains to trees, and achieves the highest average acceptance length and speedup in all four model--temperature configurations.
SwitchSD is introduced, an adaptive framework that treats copying as a latent control signal of the LLM that allows the system to dynamically switch between neural drafting and context-based copying, effectively turning copying from a noisy heuristic into a principled, model-aware decoding regime.
Roy Eisenstadt, Ido Cohen, Edo Cohen-Karlik et al.· 1 citation