Skip to content

To Copy or Not to Copy: Controlling Speculative Decoding via Intrinsic Model Signals

Sep 2026 · 1 citation · 25 references
Computer Science

TL;DR

SwitchSD is introduced, an adaptive framework that treats copying as a latent control signal of the LLM that allows the system to dynamically switch between neural drafting and context-based copying, effectively turning copying from a noisy heuristic into a principled, model-aware decoding regime.

Abstract

Speculative Decoding (SD) has significantly accelerated Large Language Model (LLM) inference, yet existing approaches face a fundamental tradeoff between two drafting strategies: neural drafting and context-based copying. Neural drafts (e.g., EAGLE3) provide robust performance across diverse text settings, while copy-based methods achieve higher speedups in copy-intensive regimes by generating candidates faster and exploiting long repetition spans for near-perfect speculation. We analyze existing copy-based methods and find that they are prone to accidental repetitions where surface-level n-gram overlap does not reflect a structural intent to copy, leading to false-positive triggers that ultimately degrade throughput. We introduce SwitchSD, an adaptive framework that treats copying as a latent control signal of the LLM. By training lightweight probes on the target model's internal representations, SwitchSD identifies genuine copy-intent with high precision (AUC>0.99). This allows the system to dynamically switch between neural drafting (e.g., EAGLE) and context-based copying. Our results across Llama and Qwen families demonstrate throughput gains of up to 15% over state-of-the-art baselines like EAGLE3, effectively turning copying from a noisy heuristic into a principled, model-aware decoding regime.

View source

Similar papers

#artificial intelligence Preprint Sep 2026

Evaluating Losslessness in Speculative Decoding Under Finite-Precision Inference

A gap between algorithmic losslessness and its implementation under finite-precision arithmetic is demonstrated and motivated, to motivate evaluating lossless speculative decoding at the level of exact generation trajectories as well as downstream task performance.

Ilya Koziev, Leonid S. Sinev, I. Oseledets · 0 citations
#artificial intelligence Preprint Sep 2026

Greedy Decoding Is Not Precision-Invariant: Cross-Precision Output Divergence in LLM Inference

An empirical error-propagation analysis is developed and finds that 22 layers of accumulated body error do not distinguish flipping from non-flipping steps; the outcome depends primarily on the top-two logit margin at the LM head relative to the directional perturbation between the top-two candidates.

Gao-Yuan Du, A. Khan, Rex Zhou et al. · 0 citations
#artificial intelligence Preprint Sep 2026

LongSpark: Efficient speculative decoding with a fixed-cost parallel drafter

Speculative decoding accelerates autoregressive inference by verifying multiple draft tokens in a single target forward pass. However, as the context grows, existing state-of-the-art drafters become increasingly expensive, eroding the very efficiency advantage they are designed to provide. We argue that this scaling is...

Hao-Yuan He, Peng-Fei Liu, Si-Shi Shen et al. · 0 citations
Preprint Aug 2026

ResiSpec: Enhancing Multi-Candidate Speculative Sampling via Residual Distribution Shaping

ResiSpec, a framework that strategically reforms the proposal distribution during verification to anchor the residual target mass within the draft model's high-confidence regions, prevents candidate obsolescence and achieves up to 1.92$\times$ speedup over state-of-the-art multi-candidate methods.

Zhi-Kai Chen, Jun Tao, Weihao Mao et al. · 1 citation
#machine learning Preprint Sep 2026

SEED: Self-Speculative Decoding via Implicit Encoder-Decoder

Self-speculative decoding accelerates large language model (LLM) inference by drafting tokens from the target model itself, but faces a sharp tradeoff between the quality and cost of the draft. Early-exit methods produce drafts cheaply by terminating computation at intermediate layers, but forgo the deeper representati...

Han-Kun Lin, Patrick Pynadath, Ruqi Zhang · 0 citations
Preprint Aug 2026

FOVEA: Focused On-Demand Visual Evidence Adaptation for Cache-Friendly Multimodal Speculative Decoding

Focused On-demand Visual Evidence Adaptation is proposed, a cache-friendly approach that builds a reusable visual memory and dynamically retrieves a bounded subset for a draft state and demonstrates that state-conditioned evidence retrieval is an effective alternative to reusing a fixed visual representation throughout...

He Zhu, Da-Yan Wu, Zihao Zhang et al. · 0 citations

Related blog posts

MIT News · Artificial Intelligence Sep 24, 2026

Estimating suicide risk from text

A new language-processing tool could help identify the highest-risk individuals from natural language, enabling swifter interventions.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.