Jul 2026
DualDecoder: Accelerate Long Context LLM Inference by Predictive Prefetch
DualDecoder is presented, a lightweight serving system for long-context LLM inference that enables efficient sparse KV cache retrieval from host memory that leverages a novel dual-token decoding pipeline that accurately identifies critical KV entries with negligible computational overhead.
Zuning Liang, Zhiyi Yao, Qi Chen et al.
· arXiv.org · 1 citation