An extended faithfulness analysis shows that the refined profiles remain largely grounded in the source preferences while preserving task-relevant personalization signals, suggesting that profile-side adaptation serves as a practical complement to universal memory construction for lifelong personalized agents.
Yuting Liu, Wei Wu, Jianzhe Zhao et al.· 0 citations
BENCHCOMPASS is introduced, a payment-domain benchmark whose construction pipeline builds scenario-grounded tasks from typed evidence packs, applies LLM-based quality checks, creates task-input attack variants, and reserves final item admission for domain experts.
Si-Jie Dong, Wei-Feng Ren, Xuan-Wei Hu et al.· 0 citations
Overall, Code2Skill transforms procedural knowledge embedded in repositories into grounded, verifiable, and transferable agent skills, providing initial evidence that the pipeline can expand with the growing volume of AI-generated software.
Yong-Qi Tong, Pan Wang, Hang Wang et al.· 1 citation
Results show that MemForest reduces memory-freshness latency while retaining strong answer quality, and introduces MemTree, a hierarchical temporal index that organizes memory as time-ordered trees and replaces global rewrites with localized dirty-path refresh.
Next-turn user reactions provide actionable local credit for improving multi-turn user-interacting agents and improves the nine-domain $\tau$-family average across three independently trained runs.
Yi-Wen Zhao, Zhihao Wen, Yuchen Mao et al.· 1 citation
DARC is proposed, a diagnosis-guided recovery harness that profiles task-family failure modes, prunes mismatched interventions from a shared recovery library, and freezes a verifier-selected success-cost policy for deployment, providing a practical route toward more reliable agents in domains where compiler-like feedba...
Pan Wang, Yihao Hu, Hang Wang et al.· 2 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.