ExplorationBench turns the wicked problem of evaluating scientific exploration into a concrete and tractable framework built on verifiable Alien Worlds, and finds that the strongest systems can acquire and apply unfamiliar rules, while performance varies substantially across trajectories and continued exploration can s...
Ming Zhang, Zhen Xiang, Pei-Zhong Gao et al.· 0 citations
Large language models (LLMs) are increasingly deployed as mobile assistants, where a key challenge is leveraging personal information scattered across multiple applications (apps) to complete user instructions. However, due to the lack of dedicated benchmarks, their capabilities remain poorly understood. To address thi...
Junjie Ye, Zhuohui Sheng, Shao-Hua Liu et al.· 0 citations
AgentGym2 is presented, a new evaluation framework with task instances grounded in real-world end-to-end working demands that measures agents'ability to execute end-to-end procedures, discover tools via exploration, compose tools for unseen tasks, and remain robust to noisy and underspecified information.
Zhiheng Xi, Dingwen Yang, Jiaqi Liu et al.· Annual Meeting of the Associ...· 1 citation
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.