ReCache: Efficient KV Cache Reuse and Compression for Tool-Augmented LLM Agents
Results show that separating reusable schema encoding from selective resource access substantially reduces agentic inference costs with limited effectiveness loss.
AI Networking Cookbook: Practical recipes for AI-assisted network automation and development
We have 2 of 7 papers
We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.
Not the right person? Other researchers publish under this name.
Results show that separating reusable schema encoding from selective resource access substantially reduces agentic inference costs with limited effectiveness loss.
WIDE is presented, the first end-to-end differentiable token-level dynamic width pruning framework designed for both prefill and decode scenarios, and a pruning--kernel co-design framework that decomposes dynamic sparsity acceleration into mask reordering, hardware-agnostic block-level skipping, and hardware-dependent intra-block skipping, enabling efficient execution across different granularities.
We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.