Preprint
Jul 2026
Less Experts, Faster Decoding: Cost-Aware Speculative Decoding for Mixture-of-Experts
Results show that accounting for expert activation cost is important for efficient speculative decoding in large-scale MoE models, and a cost-aware speculative decoding framework that incorporates predicted marginal expert activation cost into draft selection is proposed.
Jincheng Xie, Runheng Liu, Heyan Huang et al.
· 1 citation