PCR: A Prefetch-Enhanced Cache Reuse System for Low-Latency RAG Serving
Retrieval-Augmented Generation (RAG) significantly improves Large Language Models (LLMs) but introduces massive input sequences that severely bottleneck the prefill stage. While KV-cache reuse reduces redundant computation for shared document prefixes, the reusable KV working set in RAG serving can exceed GPU memory ca...