SpecEdge: Adaptive Speculative Decoding and Quantization Dynamics for Edge Small Language Models
Auto-regressive sequence generation in modern transformer language models is severely memory-bandwidth bound, yielding low arithmetic intensity on resource-constrained edge devices. While speculative decoding mitigates this bottleneck by utilizing a compact draft model to propose candidates verified concurrently by a t...