Retrieval-Augmented Generation (RAG) repeatedly prefills identical text chunks across queries, incurring redundant computations. Position-Independent Caching (PIC) mitigates it by reusing precomputed Key-Value (KV) across positions, but its efficiency is constrained by the large volume of text tokens. Rendering text ch...
Yi-Lin Liu, Rui Meng, Wang-Ze Ni et al.· 0 citations
TokenCast is proposed, which learns a composable cost representation for each execution segment, recording its own consumption and the context growth it introduces, and captures the extra input cost incurred when context from earlier segments is re-read by every later call.
Chao-Qian Ouyang, Ling Yue, Li-Bin Zheng et al.· 0 citations
This work proposes Urban-Agent, a tool-augmented agent framework for cross-system urban tasks that couples the cognitive and reasoning capabilities of a large language model with a tool-set supporting code execution, API calls, and Model Context Protocol.
Jiayu Cao, Xing-Yuan Zeng, Fei-Yue Li et al.· 0 citations
A rule-based memory framework that induces reusable logical rules from historical interactions to guide both evidence retrieval and reasoning, and constructs natural-language Horn clauses from conversations and validates them via a Rule Perplexity Consistency (RPC) mechanism.
Xing-Yuan Zeng, Zuo-Han Wu, Quanming Yao et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.