Skip to content

Evaluating Losslessness in Speculative Decoding Under Finite-Precision Inference

Sep 2026 · 0 citations · 9 references
Computer Science

TL;DR

A gap between algorithmic losslessness and its implementation under finite-precision arithmetic is demonstrated and motivated, to motivate evaluating lossless speculative decoding at the level of exact generation trajectories as well as downstream task performance.

Abstract

Lossless speculative decoding is typically defined at the algorithmic level: a speculative procedure proposes multiple tokens and a verification procedure is designed to preserve the output trajectory of an autoregressive reference model exactly. In practical neural inference, however, this guarantee is implemented using finite-precision floating-point computations, and discrete token selection can amplify small numerical differences into divergent generation trajectories. We investigate this distinction using Orthrus, a hybrid autoregressive-diffusion architecture that performs self-drafting and self-verification within a frozen autoregressive backbone, as a representative case study. Across 1,190 prompts from 12 domains, exact trajectory matching under BF16 occurs for only 45\% of the authors'checkpoint generations and 43% of those from our independently trained model. The probability of matching is strongly associated with the response-conditional perplexity of the autoregressive reference, indicating that trajectory divergence is not uniform across inputs. Despite these divergences, Orthrus does not exhibit systematic degradation on the evaluated downstream tasks. In contrast, FP32 inference yields exact trajectory matching on all evaluated prompts. These results demonstrate a gap between algorithmic losslessness and its implementation under finite-precision arithmetic, and motivate evaluating lossless speculative decoding at the level of exact generation trajectories as well as downstream task performance.

View source

Similar papers

#artificial intelligence Preprint Sep 2026

Greedy Decoding Is Not Precision-Invariant: Cross-Precision Output Divergence in LLM Inference

An empirical error-propagation analysis is developed and finds that 22 layers of accumulated body error do not distinguish flipping from non-flipping steps; the outcome depends primarily on the top-two logit margin at the LM head relative to the directional perturbation between the top-two candidates.

Gao-Yuan Du, A. Khan, Rex Zhou et al. · 0 citations
#natural language process... Preprint Aug 2026

ReTrace: Rejected-Trajectory Conditioning for Speculative Decoding

ReTrace is introduced, a rejected-trajectory conditioning method that conditions each draft block on the rejected suffix from the previous round rather than generating it from fresh mask placeholders alone, indicating that the draft model can retain useful semantic and structural information despite local token-level e...

Luxi Lin, Zhan-Peng Zeng, Shuang Peng et al. · 2 citations
Preprint Aug 2026

ResiSpec: Enhancing Multi-Candidate Speculative Sampling via Residual Distribution Shaping

ResiSpec, a framework that strategically reforms the proposal distribution during verification to anchor the residual target mass within the draft model's high-confidence regions, prevents candidate obsolescence and achieves up to 1.92$\times$ speedup over state-of-the-art multi-candidate methods.

Zhi-Kai Chen, Jun Tao, Weihao Mao et al. · 1 citation
Preprint Aug 2026

LibraSpec: Dynamic Diffusion-Based Speculative Decoding via Marginal-Gain-Driven Optimization

LibraSpec, a training-free and plug-and-play algorithm that iteratively determines the speculative length using drafter confidence scores, is developed and it is proved that LibraSpec monotonically converges toward the optimal speculative length.

Zexun Lin, Yuan Feng, Junlin Lv et al. · 0 citations
Preprint Aug 2026

DARTree: Speculative Diffusion Decoding with Autoregressive Draft Trees

DARTree is introduced, a training-free speculative decoding method that extends a pretrained AR correction head from chains to trees, and achieves the highest average acceptance length and speedup in all four model--temperature configurations.

Tian-Yi Li, Yaxin Luo, Xin-Yi Shang et al. · 4 citations · ⚡1
#natural language process... Preprint Sep 2026

To Copy or Not to Copy: Controlling Speculative Decoding via Intrinsic Model Signals

SwitchSD is introduced, an adaptive framework that treats copying as a latent control signal of the LLM that allows the system to dynamically switch between neural drafting and context-based copying, effectively turning copying from a noisy heuristic into a principled, model-aware decoding regime.

Roy Eisenstadt, Ido Cohen, Edo Cohen-Karlik et al. · 1 citation

Related blog posts

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.