GroupKV: Hierarchical KV Cache Management for Long-Context Diffusion LLM Inference
GroupKV is presented, a lightweight hierarchical KV cache management system for long-context dLLM inference that observes that under block-wise decoding, tokens within the same generation block tend to access highly overlapping and spatially concentrated context regions, making group-level sparse selection effective.