Self-evolving LLM agents promise to improve autonomously through continual interaction and learning, reducing their dependence on manually curated supervision. Realizing this promise requires not only updating the agent, but also evolving its training experience as its capabilities change. However, most existing pipeli...
Wei Yang, Shawn Li, Yuehan Qin et al.· 0 citations
CLEAR first employs a reflection agent to perform contrastive analysis over past execution trajectories and summarize useful context for each observed task to train a context augmentation model (CAM), which further optimize CAM using reinforcement learning, where the reward signal is obtained by running the task execut...
Lin-Bo Liu, Guan Wu, Han Ding et al.· arXiv.org· 0 citations
Across four RoboTwin tasks spanning different horizons and coordination patterns, Prism-GRPO improves success and quality at matched rollout budgets and reaches target success rates with up to 56% fewer rollouts.
Zeyun Deng, Yuzhe Lu, Ya-Wei Wang et al.· 1 citation
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.