The model family shows gains in held-out scientific-code repair and across selected general-purpose benchmarks in code, reasoning, and knowledge, providing evidence of positive transfer from scientific experience to broader capabilities.
He-Jia Geng, Ze-Sen Huang, Hao-Yang Li et al.· 1 citation
This work introduces and releases ScienceBuddy, an interactive scientific research workspace that brings continually improving scientific agents into researchers'everyday workflows and presents case studies of researcher interaction, harness refinement, and model learning, with the benchmark cases spanning four scienti...
Shu-Han Xue, Jian-Yuan Zhong, Ziyuan Nan et al.· 2 citations
This work forms a unified approach to capability formation and process-centered evaluation, enabling discovery behavior to be trained, improved, and measured beyond final-answer performance.
Recuris, a recursive Experiential-Working Memory architecture for long-horizon agent harnesses, in which Working Memory tracks task progress and guides skill selection from Experiential Memory, grounding skill use in current needs rather than the full history, positions recursively evolving memory as a scalable foundat...
Zhao-Chen Yu, Ying-Cheng Wu, Zhen-Fei Yin et al.· 4 citations
The PAST-Bench benchmark is introduced, a benchmark designed to isolate how persistent agents can progress from retaining experience to systematically improving through it, and Hermes+ is developed, which raises the average gain from retained experience and provides clearer pathway evidence.
Shu-Han Xue, Zixin Ding, Yi-Jun Shen et al.· 2 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.