Large language model (LLM) agents repeatedly load reusable content, such as skills, documents, and memory entries, into the current context. Re-encoding this content for every request wastes computation. Position-independent caching (PIC) alleviates this by encoding each artifact independently and reusing its key-value...
Xing-Hao Chen, Jun-Nan Dong, Cai Ke et al.· 0 citations
LGM is presented, a novel neuro-symbolic framework that shifts long-term memory disentanglement into a continuous latent space and significantly outperforms state-of-the-art baselines in capturing both explicit and implicit preferences while enabling personalized responses.
Cai Ke, Xing-Hao Chen, Xiao-Yu Shen et al.· 1 citation
This work recast CoT compression along three dimensions: importance criterion, restructuring level, and compression budget, and yields condition-aware guidelines for matching compression to deployment context.
Siyang Lyu, Zhijing Sun, Xing-Hao Chen et al.· arXiv.org· 0 citations
This paper aims to present a comprehensive overview of this emerging paradigm and establish a systematic taxonomy of latent CoT methods, categorizing them from token-wise horizontal approaches to layer-wise vertical strategies.