LLM agents rely on long-term memory to retain and reuse information when performing tasks over long horizons. Existing methods provide limited support for handling memories that become outdated as new observations or domain evidence arrive. Such outdated memories may remain semantically relevant, continue to affect dep...
Yi-Qi Wang, Jia-Qi Liu, Jia-Qi Zhang et al.· 0 citations
A language agent's execution history can exceed its context window, requiring its memory system to retrieve complete supporting evidence under a hard token budget. Evidence may span multiple execution events, yet conventional retrievers use fixed token windows and fixed-k metrics that reward individual fragments withou...
Yi-Qi Wang, Jin-Qian Ju, Jia-Qi Zhang et al.· 0 citations
It is shown that this loop can sustain a lower-return policy even when representation fitting is globally optimal on data selected by the agent, which then uses the resulting returns to guide its next choices.
Evaluating claim admission in shared agent memory is challenging because repeated claims may be mistaken for independent evidence. An agent may copy or paraphrase a retrieved belief, while admitting a false claim exposes subsequent agents to it. To study this problem, we introduce the Correlated Promotion Benchmark (CP...
Xiao-Yang Li, Yi-Qi Wang, Chen-Cheng Zhu et al.· 0 citations
IndustrialVLA-Bench is presented, an evidence-aware evaluation of six released VLA and WAM systems under a unified reporting schema that evaluates clean capability on LIBERO, non-language robustness on LIBERO-Plus, instruction sensitivity on LIBERO-Para, and observed execution cost.
Yi-Qi Wang, Zhi-Feng Rao, Jia-Qi Zhang et al.· 0 citations
The results do not imply uniformly better trace reconstruction, but show that dependency-guided rollback repair provides a strong recovery--cost trade-off while repairing faulty memory state and preserving benign memory.
Cailing Yu, Yiqi Wang, Jiaqi Zhang et al.· 4 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.