Skip to content

Author

Zirou Liu

We have 2 of 2 papers

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

Jul 2026

How Benchmarks Mis-Score Computer-Use Agents

Computer-use agents (CUA) are being deployed to browse the web and operate desktop software, yet their benchmark scores are still commonly produced by brittle scripted oracles. A score is the output of a pipeline in which tasks can be stale, trajectories can omit decisive visual evidence, evaluators can reject valid alternatives, and aggregate reports can hide the cause of failure. We organize these problems into a reliability framework spanning task construction, trajectory observation, scoring, and reporting. We then audit 150 public failure-scored trajectories from five web, enterprise-workflow, and desktop-control benchmarks, find that 15.3\% of FAIL verdicts are wrong: 10.7\% are evaluator false negatives and 4.7\% are broken tasks. For genuine failures, a three-tier diagnostic taxonomy shows that verification/feedback and planning failures dominate execution/grounding errors, while a single scalar success rate can not explain. We connect these findings to newer long-horizon CUA benchmarks and derive stage-specific design rules for CUA evaluation.

Zi-Han Dong, Zhiyuan Ma, Zekun Wang et al. · 2 citations
Review Aug 2026

Agentic Artifact Creation: Systems, Evaluation, Principles, and Opportunities

This survey examines agentic artifact creation, which is defined as stateful construction in which an AI system materially constructs or revises a deliverable and intermediate observations redirect later work, and formulate principles for keeping commitments and responsibility explicit, turning feedback into targeted repair, and revalidating affected state after change.

Tianfu Wang, Zhezheng Hao, Xinchi Xia et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.