Preprint
Aug 2026
ReCache: Efficient KV Cache Reuse and Compression for Tool-Augmented LLM Agents
Results show that separating reusable schema encoding from selective resource access substantially reduces agentic inference costs with limited effectiveness loss.
Yichu Fang, Sitong Wei, Haozhe Hu et al.
· 1 citation