CatKV: Accelerating Position-Independent Caching with Context-Adaptive KV Cache Compression
Context caching is widely adopted to reduce the dominant prefill cost for long prompts across various large language model (LLM) applications. Recently, position-independent caching (PIC) has emerged to improve KV cache reuse beyond strict prefix matching. However, the required KV caches often reside in a persistent re...