Many real-world tasks (e.g., office workflows, scientific experimentation) require LLM agents to interact repeatedly with their environments for context-dependent operations. However, such environments are often not agent-ready. First, information is often scattered and fragmented across the environment. Second, releva...
Yu-Kai Wu, Yuan-Jing Yang, Leon Zhou et al.· 0 citations
Complex questions often require multi-hop reasoning that connects facts distributed across sources or distant regions of a long context through intermediate steps. Benchmarks commonly evaluate this ability with questions built around predefined reasoning chains, treating a correct answer as evidence that the intended c...
Ji-Hua Tao, Xiao-Kun Yuan, Yao-Ming Li et al.· 0 citations
Human evaluation provides a direct measure of the quality of LLM-generated web applications. However, fitting human judgments through automated evaluation remains challenging. Static benchmarks can credit functionality that exists in source code but is unreachable at runtime. Interactive benchmarks exercise the applica...
Chen-Xu Liu, Zi-Lu Zou, Pei-Zhong Gao et al.· 0 citations
ExplorationBench turns the wicked problem of evaluating scientific exploration into a concrete and tractable framework built on verifiable Alien Worlds, and finds that the strongest systems can acquire and apply unfamiliar rules, while performance varies substantially across trajectories and continued exploration can s...
Ming Zhang, Zhen Xiang, Pei-Zhong Gao et al.· 0 citations
Env-Rethink is proposed, a system with 27B post-trained model that supports three main capabilities that adaptively builds Collection Maps and Event Logs to supplement necessary context and evolves environments through virtual event histories that alter environmental states and evidence relationships.
Yu-Kai Wu, Yuan-Jing Yang, Leon Zhou et al.· 0 citations
This work proposes DualViewEval, an agent benchmark compression method that jointly exploits outcome and process relations to learn an exact-size miniset and predict the full-benchmark scores, and achieves the best results in all datasets.
Xin-Shuai Guo, Jun-Jie Wu, Dolly Deng et al.· 0 citations
Human evaluation provides a direct measure of the quality of LLM-generated web applications. However, fitting human judgments through automated evaluation remains challenging. Static benchmarks can credit functionality that exists in source code but is unreachable at runtime. Interactive benchmarks exercise the applica...
Chen-Xu Liu, Zi-Lu Zou, Pei-Zhong Gao et al.· 0 citations
Agent benchmarks increasingly evaluate repository editing, web research, terminal use, and long-horizon interaction. Their scores support capability claims only when the evaluation protocol keeps the intended capability necessary for success. Recent reward-hacking benchmarks and system reports show that agents can inst...
This work analyzes real chatbot failures to identify six recurring mechanisms and defines six controlled evaluation modes in Hy-MultiTurn, a Chinese benchmark for deep multi-turn dialogue understanding, which shows that Hy-MultiTurn is broadly challenging.
Eileen Ye, Ji-Hua Tao, Yao-Ming Li et al.· arXiv.org· 0 citations
E-Bench is introduced, a fully synthetic benchmark with 323 state-changing tasks across three product domains: Honor of Kings, QQ Music, and Tencent Meeting, and it shows that multi-step tool use remains challenging: Pass^3 stays below 60% for the strongest models, and even with code execution in the E-Bench-Code exten...
Weihuang Zheng, Tianyuan Zou, Eileen Ye et al.· arXiv.org· 1 citation
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.