RR-Evict: Fine-Grained Prefix Cache Eviction beyond LRU for Agentic LLM Serving
LLM-based agents execute long-horizon tasks through repeated model calls interleaved with tool execution and user interaction. As each call extends the history accumulated in previous turns, prefix caching avoids repeated prefill of the agent's entire context. However, the aggregate cache footprint grows with context l...