Skip to content

GroupKV: Hierarchical KV Cache Management for Long-Context Diffusion LLM Inference

Jul 2026 · Proceedings of the International Conference on Parallel Processing · pp. 553-563 · 0 citations · 44 references
Computer Science

Abstract

Diffusion large language models (dLLMs) are emerging as a promising generative paradigm that complements autoregressive decoding. In long-context settings, KV cache bloat and offloading transfer overhead have become primary bottlenecks in inference systems. Meanwhile, the periodic full-sequence recomputation and localized token updates in dLLMs make the KV lifecycle substantially more dynamic, complicating cache management and prefetch scheduling while making heavyweight token-level indexing or clustering schemes harder to amortize effectively during decoding. To address these challenges, we present GroupKV, a lightweight hierarchical KV cache management system for long-context dLLM inference. We observe that under block-wise decoding, tokens within the same generation block tend to access highly overlapping and spatially concentrated context regions, making group-level sparse selection effective. Building on this observation, GroupKV partitions the context into contiguous groups and performs coarse-to-fine sparse selection. GroupKV further exploits cross-layer consistency to enable predictive prefetching, and incorporates a staleness correction mechanism to maintain cache coherence under dynamic KV updates. Additionally, GroupKV adopts streaming prefill to reduce peak memory consumption during prefilling. Experiments show that GroupKV extends the maximum serviceable context length by up to 48.00 × under constrained GPU memory, improves end-to-end inference performance by up to 3.73 × in offload-based long-context settings, and maintains competitive task accuracy.

Read PDF

Similar papers

Book Open access Aug 2026

Turbo: Efficiently Serving Long-Context Large Language Models with In-Network Aggregation

This work proposes Turbo, a first-of-its-kind in-network aggregation system that accelerates long-context inference by offloading query broadcast and attention aggregation to switches and introduces a rolling forward scheme that propagates states to enable cross-stage updates.

Ying Wan, Yuchen Xu, Chuwen Zhang et al. · 0 citations
Preprint Aug 2026

OasisKV: Scaling In-Decode KV Cache Beyond HBM with Lookahead Sparse Prefetching

OasisKV is presented, a memory-centric LLM inference system design that alleviates HBM capacity pressure by decoupling full KV-cache storage from HBM during LLM decoding and observes that future important tokens can be predicted accurately in advance using lookahead tokens drafted by speculative decoding (SD).

Can Xiao, Sukmin Cho, J. We et al. · 0 citations
#machine learning Preprint Sep 2026

TierKV: Long-Context On-Device LLMs via Predictive Multi-Tier KV Caching

TierKV is presented, a mobile LLM inference framework built on Predictive Multi-Tier Cache Optimization (PMCO), which improves prefill throughput by up to 17.6x over existing mobile LLM frameworks, reduces RAM-resident KV cache by 12.5-34%, thereby enabling substantially longer contexts under the same memory budget, wh...

Zhi-Hao Shu, Md Musfiqur Rahman Sanim, Jie Hu et al. · 0 citations
Preprint Aug 2026

HiSparse: Scaling Sparse-Attention Decoding with Hierarchical KV Cache Management

HiSparse is merged into upstream SGLang and evaluated across three sparse-attention families (DSA, NSA, and Quest) on H200, B200, and GH200 platforms: it improves peak generation throughput by up to 4.7x on long-context workloads while preserving comparable per-token latency and reducing time-to-first-token at high loa...

Zhiqiang Xie, Zhangheng Huang, Ting-Jun Huang et al. · 4 citations · ⚡1
#artificial intelligence Preprint Sep 2026

HeadWiseKV: Budgeted Per-Head Cache Residency for Hybrid Long-Context Language Models

This work introduces HeadWiseKV, a training-free framework that compresses the residual global KV caches of hybrid language models while preserving their native local, recurrent, and linear paths and proposes SeqCalib as the core policy-generation algorithm in HeadWiseKV.

Ren-Jie Xie, Jun-Cheng Yang, Ao-Ting Hu et al. · 0 citations
#artificial intelligence Preprint Sep 2026

CompKV: Compensation-Aware KV Selection for Long-Context LLM Inference

CompKV is introduced, the first compensation-aware sparse attention framework that divides tokens into blocks and explicitly optimizes selection for the downstream compensation mechanism, and shows that the residual left by block-level mean compensation is governed by both block attention mass and within-block logit va...

Zheng-Hong Huang, Rui-Zhe Yao, Dan-Yi Liu et al. · 1 citation · ⚡1

Related blog posts

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.