Skip to content

Author

Maxm Pan

10 papers indexed here

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

Preprint Sep 2026

Breaking the Environment Wall: A Unified Framework for Preparing and Evolving Agent-Native Environments

Many real-world tasks (e.g., office workflows, scientific experimentation) require LLM agents to interact repeatedly with their environments for context-dependent operations. However, such environments are often not agent-ready. First, information is often scattered and fragmented across the environment. Second, releva...

Yu-Kai Wu, Yuan-Jing Yang, Leon Zhou et al. · 0 citations
#artificial intelligence Preprint Sep 2026

REALHOP: Rethinking Multi-Hop Reasoning Evaluation via Behavioral Auditing

Complex questions often require multi-hop reasoning that connects facts distributed across sources or distant regions of a long context through intermediate steps. Benchmarks commonly evaluate this ability with questions built around predefined reasoning chains, treating a correct answer as evidence that the intended c...

Ji-Hua Tao, Xiao-Kun Yuan, Yao-Ming Li et al. · 0 citations
Preprint Sep 2026

WebCraftBench: Evaluating Web Application Generation from a Software Testing Perspective

Human evaluation provides a direct measure of the quality of LLM-generated web applications. However, fitting human judgments through automated evaluation remains challenging. Static benchmarks can credit functionality that exists in source code but is unreachable at runtime. Interactive benchmarks exercise the applica...

Chen-Xu Liu, Zi-Lu Zou, Pei-Zhong Gao et al. · 0 citations
#artificial intelligence Preprint Sep 2026

ExplorationBench: Measuring AI Systems'Exploration in Verifiable Alien Worlds

ExplorationBench turns the wicked problem of evaluating scientific exploration into a concrete and tractable framework built on verifiable Alien Worlds, and finds that the strongest systems can acquire and apply unfamiliar rules, while performance varies substantially across trajectories and continued exploration can s...

Ming Zhang, Zhen Xiang, Pei-Zhong Gao et al. · 0 citations
#artificial intelligence Preprint Sep 2026

Breaking the Environment Wall: Evolving LLM Agent Environments for Recursive Self-Improvement

Env-Rethink is proposed, a system with 27B post-trained model that supports three main capabilities that adaptively builds Collection Maps and Event Logs to supplement necessary context and evolves environments through virtual event histories that alter environmental states and evidence relationships.

Yu-Kai Wu, Yuan-Jing Yang, Leon Zhou et al. · 0 citations
#artificial intelligence Preprint Sep 2026

IWC-Bench: Evaluating Web Application Generation from a Software Testing Perspective

Human evaluation provides a direct measure of the quality of LLM-generated web applications. However, fitting human judgments through automated evaluation remains challenging. Static benchmarks can credit functionality that exists in source code but is unreachable at runtime. Interactive benchmarks exercise the applica...

Chen-Xu Liu, Zi-Lu Zou, Pei-Zhong Gao et al. · 0 citations
Jul 2026

Do Agent Benchmarks Measure Capability? Protocol Validity in the Age of Agentic AI

Agent benchmarks increasingly evaluate repository editing, web research, terminal use, and long-horizon interaction. Their scores support capability claims only when the evaluation protocol keeps the intended capability necessary for success. Recent reward-hacking benchmarks and system reports show that agents can inst...

Jia-Qi Shao, Hanck Chen, Wei Zhang et al. · 4 citations
Jul 2026

Hy-MultiTurn: A Six-Dimensional Benchmark for Deep Multi-Turn Dialogue Understanding

This work analyzes real chatbot failures to identify six recurring mechanisms and defines six controlled evaluation modes in Hy-MultiTurn, a Chinese benchmark for deep multi-turn dialogue understanding, which shows that Hy-MultiTurn is broadly challenging.

Eileen Ye, Ji-Hua Tao, Yao-Ming Li et al. · 0 citations
Jul 2026

E-Bench: Benchmarking Multi-Step Tool-Use Agents in Real-World Product Scenarios

E-Bench is introduced, a fully synthetic benchmark with 323 state-changing tasks across three product domains: Honor of Kings, QQ Music, and Tencent Meeting, and it shows that multi-step tool use remains challenging: Pass^3 stays below 60% for the strongest models, and even with code execution in the E-Bench-Code exten...

Weihuang Zheng, Tianyuan Zou, Eileen Ye et al. · 1 citation

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.