Results show that long-horizon reflective data is an effective route toward self-improving agents, and synthesize long-horizon improvement trajectories from machine learning and algorithmic programming tasks, two domains that offer verifiable feedback and reward sustained iteration.
Hong-Jin Qian, Chao-Fan Li, Kun Luo et al.· 0 citations
Memory is critical for AI agents. Many existing agent-memory systems follow an Ahead-of-Time (AOT) design, constructing memory before a specific request arrives. While this reduces online serving cost, such request-agnostic memory construction can discard fine-grained information that later becomes important. To addres...
Bing-Yu Yan, Chao-Fan Li, Hong-Jin Qian et al.· 0 citations
This work introduces AREX, a family of Recursively Self-Improving (RSI) deep research agents that substantially outperforms comparable-scale baselines and remains competitive with models using substantially more activated parameters.
Shuqi Lu, Chaofan Li, Kun Luo et al.· arXiv.org· 2 citations· ⚡1
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.