Reimagining LLM Inference Infrastructure with Memory-Centric KV Cache Servers
It is argued that a KV cache server — disaggregated, CXL-attached memory device(s) with optional near-memory compute—is the correct infrastructure primitive for the next generation of AI data centers.