LLM agents increasingly maintain personal memory across sessions, but it can conflict. Preferences depend on context, behavior evolves, and sources can conflict. When a query lacks context, time, or source authority to interpret conflict, treating one memory as definitive converts unresolved conflict into an unjustifie...
Lulu Yang, Shusheng Xu, Zhuo-Ran Li et al.· 1 citation
Multi-agent reinforcement learning (MARL) provides a powerful framework for learning coordinated behaviors through interactions with the environment. Developing MARL policies requires balancing expressive modeling of complex and multimodal action distributions with efficient training and execution. Generative policies,...
Zhuo-Ran Li, Yun-Zhan Li, Xun Wang et al.· 0 citations
FlexLoop is proposed, a novel post-training framework that converts pretrained fixed-depth looped policies into depth-elastic policies that supports reliable inference across recurrent depths and enables state-wise adaptive inference through recurrent-depth consistency.
Xun Wang, Rui-Shuo Chen, Yu Chen et al.· 0 citations
The proposed Module Level Reward Evolution Framework integrates three mechanisms: reflection-based refinement, hybrid credit assignment, and a merge strategy with rollback, which together improve the effectiveness and robustness of reward optimization.
Chenglin Liu, Xun Wang, Ruishuo Chen et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.