#natural language process...
May 2026
GRKV: Global Regression for Training-Free KV Cache Compression in Long-Context LLMs
GRKV (Global Regression for KV Cache), a training-free KV-cache merging method that directly minimizes the discrepancy between compressed-cache and full-cache attention outputs, is proposed.
Junjie Peng, You Wu, Haoyi Wu et al.
· arXiv.org · 2 citations