Jul 2026· Proceedings of the International Conference on Parallel Processing· pp. 553-563· 0 citations· 44 references
Computer Science
Abstract
Diffusion large language models (dLLMs) are emerging as a promising generative paradigm that complements autoregressive decoding. In long-context settings, KV cache bloat and offloading transfer overhead have become primary bottlenecks in inference systems. Meanwhile, the periodic full-sequence recomputation and localized token updates in dLLMs make the KV lifecycle substantially more dynamic, complicating cache management and prefetch scheduling while making heavyweight token-level indexing or clustering schemes harder to amortize effectively during decoding. To address these challenges, we present GroupKV, a lightweight hierarchical KV cache management system for long-context dLLM inference. We observe that under block-wise decoding, tokens within the same generation block tend to access highly overlapping and spatially concentrated context regions, making group-level sparse selection effective. Building on this observation, GroupKV partitions the context into contiguous groups and performs coarse-to-fine sparse selection. GroupKV further exploits cross-layer consistency to enable predictive prefetching, and incorporates a staleness correction mechanism to maintain cache coherence under dynamic KV updates. Additionally, GroupKV adopts streaming prefill to reduce peak memory consumption during prefilling. Experiments show that GroupKV extends the maximum serviceable context length by up to 48.00 × under constrained GPU memory, improves end-to-end inference performance by up to 3.73 × in offload-based long-context settings, and maintains competitive task accuracy.
This work proposes Turbo, a first-of-its-kind in-network aggregation system that accelerates long-context inference by offloading query broadcast and attention aggregation to switches and introduces a rolling forward scheme that propagates states to enable cross-stage updates.
Ying Wan, Yuchen Xu, Chuwen Zhang et al.· Conference on Applications,...· 0 citations
OasisKV is presented, a memory-centric LLM inference system design that alleviates HBM capacity pressure by decoupling full KV-cache storage from HBM during LLM decoding and observes that future important tokens can be predicted accurately in advance using lookahead tokens drafted by speculative decoding (SD).
TierKV is presented, a mobile LLM inference framework built on Predictive Multi-Tier Cache Optimization (PMCO), which improves prefill throughput by up to 17.6x over existing mobile LLM frameworks, reduces RAM-resident KV cache by 12.5-34%, thereby enabling substantially longer contexts under the same memory budget, wh...
Zhi-Hao Shu, Md Musfiqur Rahman Sanim, Jie Hu et al.· 0 citations
HiSparse is merged into upstream SGLang and evaluated across three sparse-attention families (DSA, NSA, and Quest) on H200, B200, and GH200 platforms: it improves peak generation throughput by up to 4.7x on long-context workloads while preserving comparable per-token latency and reducing time-to-first-token at high loa...
This work introduces HeadWiseKV, a training-free framework that compresses the residual global KV caches of hybrid language models while preserving their native local, recurrent, and linear paths and proposes SeqCalib as the core policy-generation algorithm in HeadWiseKV.
Ren-Jie Xie, Jun-Cheng Yang, Ao-Ting Hu et al.· 0 citations
CompKV is introduced, the first compensation-aware sparse attention framework that divides tokens into blocks and explicitly optimizes selection for the downstream compensation mechanism, and shows that the residual left by block-level mean compensation is governed by both block attention mass and within-block logit va...
Zheng-Hong Huang, Rui-Zhe Yao, Dan-Yi Liu et al.· 1 citation· ⚡1