Skip to content

Author

Qi Zhang

We have 11 of 24 papers

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

#artificial intelligence Preprint Sep 2026

EvoIn: Bridging Evolution and Internalization for Agent Fine-Tuning

EvoIn is an agent fine-tuning framework that bridges evolution and internalization, and consistently enables agents to learn stronger decision-making procedures, raising the pass rate by 10.9 points in-domain and by 9.2 points out-of-domain.

Shi-Han Dou, Shao-Hua Liu, Zhong-Hang Lu et al. · 0 citations
#artificial intelligence Preprint Sep 2026

SciHorizon-eLab: An Agentic Protocol-to-Task Compiler for Scalable Benchmarking of Scientific Embodied Agents

Embodied agents offer a promising route to automating scientific experimentation, yet their progress is constrained by the lack of reliable and systematic evaluation environments. Existing simulation-based laboratory benchmarks rely heavily on manual task engineering, making it challenging to systematically compile div...

Mao-Kai Qin, Chuan Qin, Qi Zhang et al. · 0 citations
#artificial intelligence Preprint Sep 2026

ExplorationBench: Measuring AI Systems'Exploration in Verifiable Alien Worlds

ExplorationBench turns the wicked problem of evaluating scientific exploration into a concrete and tractable framework built on verifiable Alien Worlds, and finds that the strongest systems can acquire and apply unfamiliar rules, while performance varies substantially across trajectories and continued exploration can s...

Ming Zhang, Zhen Xiang, Pei-Zhong Gao et al. · 0 citations
#artificial intelligence Preprint Sep 2026

Strategy Accumulation and Guided Execution for Automated LLM Fine-Tuning

Producing task-specific large language models requires discovering effective training strategies through experimentation. Automated fine-tuning systems have made this experimentation feasible with far less manual effort. However, these systems are stateless: each search discards its discovered strategies, dataset insig...

Hao-Ran Zhao, Wei Du, Dingwen Yang et al. · 0 citations
Preprint Aug 2026

SPIEval: Evaluating Large Language Models as Mobile Assistants over Scattered Personal Information

Large language models (LLMs) are increasingly deployed as mobile assistants, where a key challenge is leveraging personal information scattered across multiple applications (apps) to complete user instructions. However, due to the lack of dedicated benchmarks, their capabilities remain poorly understood. To address thi...

Junjie Ye, Zhuohui Sheng, Shao-Hua Liu et al. · 0 citations
#artificial intelligence Preprint Sep 2026

SWE-Bench Pro Verified: A Reliable Benchmark for Software Engineering Agents

SWE-Bench Pro Verified offers a more trustworthy benchmark for assessing software engineering agents, which combines anti-hacking safeguards that eliminate major leakage channels without disrupting normal agent functionality, with task refinement that minimally corrects inconsistencies within flawed instances.

Pujun Zheng, Zi-Xin Shang, Shufan Jiang et al. · 1 citation
Jul 2026

Beacon: Knowing When and How to Perform Agentic Visual Reasoning

The fundamental goal of agentic visual reasoning is to improve the success rate of multimodal large language models (MLLMs) on complex tasks, rather than merely equipping them with a sophisticated yet inefficient reasoning paradigm. In this work, we rethink agentic visual reasoning through two key dimensions of tool us...

Qixun Wang, Yang Shi, Le-Tian Cheng et al. · 0 citations

JFTA-Bench: Evaluate LLM's Ability of Tracking and Analyzing Malfunctions Using Fault Trees

A novel textual representation of fault trees is proposed, and a benchmark for multi-turn dialogue systems that emphasizes robust interaction in complex environments is constructed, evaluating a model's ability to assist in malfunction localization.

Yuhui Wang, Zhi-Xiong Yang, Ming Zhang et al. · 0 citations
Jul 2026

AgentCompass: A Unified Evaluation Infrastructure for Agent Capabilities

AgentCompass is introduced, an open-source, lightweight, and extensible infrastructure for evaluating LLM-based agents that organizes the evaluation process around three independent components, thereby enabling flexible configurations without requiring the reimplementation of complex execution logic.

Zichen Ding, Jiaye Ge, Shufan Jiang et al. · 2 citations
Preprint Aug 2026

CAFE: Self-Improving Search Agents Need Co-Evolving Feedback

CAFE (Coupled Agent--Feedback Evolution), a framework in which a shared-parameter model alternates between search-agent and critic roles, is introduced, suggesting that a self-improving search agent needs feedback that co-evolves with the policy it guides.

Bo-Yang Liu, Sen-Jie Jin, Pei-Xin Wang et al. · 1 citation
Jul 2026

HiSkill: Empowering LLM Agents with Hierarchical Skill Graphs

Experiments show that HiSkill outperforms state-of-the-art baselines while reducing inference token consumption, demonstrating the effectiveness of bridging high-level skills and executable action grounding through a hierarchical skill graph.

Yu Hao, Jinxuan Cai, Qi Zhang et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.