Reported Speculative decoding (SD) speedups are difficult to compare because methods are commonly evaluated with different serving runtimes and configurations. We present SpecLLM, a pluggable evaluation framework that applies common scheduling, batching, KV-cache, CUDA Graph, and execution-backend policies across methods while preserving method-required proposal and verification paths. We implement eight Speculative Decoding (SD) methods and evaluate them across batch sizes, controlled online loads, and two model scales. We decompose throughput speedup as S=τ/c, where τ is the average number of committed tokens per SD step and c is the dimensionless execution cost of that step relative to a matched autoregressive decode step. The measured c is conditional on the method–runtime interaction, hardware, workload, baseline, and configuration. Under the evaluated conditions, method rankings change with batch size and online load, and a longer commit length does not necessarily yield lower latency when execution cost or queueing increases. SpecLLM therefore provides controlled within-runtime comparability rather than runtime-independent fairness; SD results should report (τ,c), latency, feasibility, and the conditions under which they were measured.
Sungkyun Kim, Jaemin Kim, Yeongpil Cho et al.· Electronics· 0 citations
NeuRIT is proposed, a Neuron-guided Robust Instruction-Tuning framework built on a localization-first perspective that mines context-aware neurons associated with relevant and irrelevant context processing, and uses them as anchors to selectively adapt both the identified neuron groups and the layers in which they concentrate.
Jae Lee, Jaemin Kim, Sumyeong Ahn et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.