Memory self-evolution uses task feedback to iteratively improve executable memory programs that store and retrieve information from past interactions. Existing approaches typically adopt holistic evolution, deriving revision directions from mixed feedback and judging progress by overall performance. This can obscure op...
Yao-Qi Chen, Yu-Ru Feng, Qianxi Zhang et al.· 0 citations
PEARL is an asynchronous agentic RL system that coordinates external resource elasticity, temporary reuse of idle training GPUs, and adaptive PD execution, and maintains a unified GPU--worker--role state and uses runtime profiles to predict rollout batch completion time.
Ji-Aan Zhu, Wei Gao, You-Hui Bai et al.· 0 citations
DeaMoE is proposed, a decoding-efficient MoE architecture, in which the experts are grouped into several departments, and the experts belonging to the same department share most parameters since they come from the same professional field, and additionally each expert contains a few private parameters to reflect its uni...
Ze-Wen Jin, Sheng-Yu Fu, Zeping Duan et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.