Skip to content
Book Open access

LLM KV-cache: To Restore or To Recompute, That Is the Question

Sep 2026 · Proceedings of the 18th ACM Workshop on Hot Topics in Storage and File Systems · 0 citations · 11 references

TL;DR

The challenges of the restoration-recomputation trade-off are investigated and its impact on inference performance when left unaddressed, and an I/O-aware KV-cache management policy is presented that dynamically navigates this trade-off.

Abstract

Key-Value (KV)-caches are essential to modern Large Language Model (LLM) inference, transforming inference from a computationally prohibitive task into a practical one. However, the limited capacity of GPU High Bandwidth Memory (HBM) is often insufficient to retain KV-cache data for all active requests, especially when large models and long contexts are served. To address this limitation, modern inference systems support offloading KV-cache data from HBM to CPU DRAM, local, or remote storage. A known issue is the trade-off between restoring offloaded KV-cache data versus recomputing; the right policy is non-trivial and depends on dynamic factors such as storage bandwidth, interconnect performance, workload characteristics, context length, and service-level objectives (SLOs). In this paper, we first investigate the challenges of the restoration-recomputation trade-off and quantify its impact on inference performance when left unaddressed. We then present an I/O-aware KV-cache management policy that dynamically navigates this trade-off. Our approach maximizes inference throughput while satisfying service-level objectives by either restoring or recomputing KV-cache blocks based on the performance characteristics of the GPU and storage tiers. Initial theoretical results show a 10 × improvement in performance over simplistic static policies while maintaining SLO compliance.

Read PDF

Similar papers

Understanding and Optimizing KV-cache Management for Long-Context LLM Inference A

This model reveals one key opportunity: dividing a restore request proportionally between the storage path and the GPU can improve inference performance while still meeting SLOs, and reduces the KV-cache storage stack to a performance model based on per-tier capacity, per-tier and interconnect bandwidth, and GPU arithm...

Unknown authors · 0 citations
Preprint Aug 2026

OasisKV: Scaling In-Decode KV Cache Beyond HBM with Lookahead Sparse Prefetching

OasisKV is presented, a memory-centric LLM inference system design that alleviates HBM capacity pressure by decoupling full KV-cache storage from HBM during LLM decoding and observes that future important tokens can be predicted accurately in advance using lookahead tokens drafted by speculative decoding (SD).

Can Xiao, Sukmin Cho, J. We et al. · 0 citations
#machine learning Preprint Sep 2026

MetaKV: Adaptive KV Cache Compression for Constrained LLM Inference

Key--value (KV) cache compression is an effective way to reduce the memory overhead of large language model (LLM) inference, particularly for long-context workloads. However, existing compression methods make different trade-offs among accuracy, inference latency, and peak KV cache memory utilization, making a single f...

Michael Wang, Keith Li, Roozbeh Bostandoost · 0 citations
Conference Aug 2026

Characterizing Predictability–Latency Trade-offs of KV-Cache SSD Offloading in LMCache for LLM Serving Systems

KV-cache offload is widely used to stretch GPU memory for LLM serving, but its storage behavior has not been characterized at the block-device level. In this paper, we study LMCache through realworld multi-session workloads that span same/different context $\times$ same/different prompt, using over 100 stateless reques...

Ying He, Dingsen Shi, Yanbo Dai et al. · 0 citations
Book Open access Sep 2026

To Keep or Not to Keep: Learning KV Cache Retention in Disaggregated LLM Serving Systems

Disaggregated LLM serving separates prefill and decode into distinct node pools, interposing a network fabric between the moment a key-value (KV) cache is computed and the moment it is consumed. This architectural shift invalidates a core assumption of classical cache policies: that the cost of a miss is simply recompu...

Dong Liu, Yan-Xuan Yu, Eric Jiang et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.