Large language model (LLM) agents repeatedly load reusable content, such as skills, documents, and memory entries, into the current context. Re-encoding this content for every request wastes computation. Position-independent caching (PIC) alleviates this by encoding each artifact independently and reusing its key-value...
Xing-Hao Chen, Jun-Nan Dong, Cai Ke et al.· 0 citations
Parallel speculative drafting generates multiple candidates in one backbone pass, but independent token selection can produce inconsistent continuations that shorten the accepted prefix. Existing methods mostly leave conditional decoding to a lightweight module after the backbone, which limits the flow of predecessor i...
Hao-Hui Zhang, Ke-Yu Chen, Hao-Cheng Sun et al.· 0 citations
LGM is presented, a novel neuro-symbolic framework that shifts long-term memory disentanglement into a continuous latent space and significantly outperforms state-of-the-art baselines in capturing both explicit and implicit preferences while enabling personalized responses.
Cai Ke, Xing-Hao Chen, Xiao-Yu Shen et al.· 1 citation
Experiments reveal that models with similar end-to-end accuracy can exhibit markedly different agentic capability profiles, demonstrating that process-level evaluation is crucial for interpreting the true potential of LLMs and guiding the development of next-generation mathematical agents.
Jiayi Kuang, Ying-Hui Li, Yun-Ze Song et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.