EvoIn is an agent fine-tuning framework that bridges evolution and internalization, and consistently enables agents to learn stronger decision-making procedures, raising the pass rate by 10.9 points in-domain and by 9.2 points out-of-domain.
Shi-Han Dou, Shao-Hua Liu, Zhong-Hang Lu et al.· 0 citations
Embodied agents offer a promising route to automating scientific experimentation, yet their progress is constrained by the lack of reliable and systematic evaluation environments. Existing simulation-based laboratory benchmarks rely heavily on manual task engineering, making it challenging to systematically compile div...
Mao-Kai Qin, Chuan Qin, Qi Zhang et al.· 0 citations
ExplorationBench turns the wicked problem of evaluating scientific exploration into a concrete and tractable framework built on verifiable Alien Worlds, and finds that the strongest systems can acquire and apply unfamiliar rules, while performance varies substantially across trajectories and continued exploration can s...
Ming Zhang, Zhen Xiang, Pei-Zhong Gao et al.· 0 citations
Producing task-specific large language models requires discovering effective training strategies through experimentation. Automated fine-tuning systems have made this experimentation feasible with far less manual effort. However, these systems are stateless: each search discards its discovered strategies, dataset insig...
Hao-Ran Zhao, Wei Du, Dingwen Yang et al.· 0 citations
Large language models (LLMs) are increasingly deployed as mobile assistants, where a key challenge is leveraging personal information scattered across multiple applications (apps) to complete user instructions. However, due to the lack of dedicated benchmarks, their capabilities remain poorly understood. To address thi...
Junjie Ye, Zhuohui Sheng, Shao-Hua Liu et al.· 0 citations
SWE-Bench Pro Verified offers a more trustworthy benchmark for assessing software engineering agents, which combines anti-hacking safeguards that eliminate major leakage channels without disrupting normal agent functionality, with task refinement that minimally corrects inconsistencies within flawed instances.
Pujun Zheng, Zi-Xin Shang, Shufan Jiang et al.· 1 citation
The fundamental goal of agentic visual reasoning is to improve the success rate of multimodal large language models (MLLMs) on complex tasks, rather than merely equipping them with a sophisticated yet inefficient reasoning paradigm. In this work, we rethink agentic visual reasoning through two key dimensions of tool us...
Qixun Wang, Yang Shi, Le-Tian Cheng et al.· arXiv.org· 0 citations
A novel textual representation of fault trees is proposed, and a benchmark for multi-turn dialogue systems that emphasizes robust interaction in complex environments is constructed, evaluating a model's ability to assist in malfunction localization.
Yuhui Wang, Zhi-Xiong Yang, Ming Zhang et al.· arXiv.org· 0 citations
AgentCompass is introduced, an open-source, lightweight, and extensible infrastructure for evaluating LLM-based agents that organizes the evaluation process around three independent components, thereby enabling flexible configurations without requiring the reimplementation of complex execution logic.
CAFE (Coupled Agent--Feedback Evolution), a framework in which a shared-parameter model alternates between search-agent and critic roles, is introduced, suggesting that a self-improving search agent needs feedback that co-evolves with the policy it guides.
Bo-Yang Liu, Sen-Jie Jin, Pei-Xin Wang et al.· 1 citation
Experiments show that HiSkill outperforms state-of-the-art baselines while reducing inference token consumption, demonstrating the effectiveness of bridging high-level skills and executable action grounding through a hierarchical skill graph.