Skip to content
Preprint

AgentWorld: Benchmarking Long-Horizon Collaboration of Multi-agent LLMs

Sep 2026 · 0 citations · 37 references
Computer Science

TL;DR

To quantify collaboration effectiveness in addition to conventional binary task success, Causal Collaboration Effectiveness (CCE) is proposed, a graph-based metric that traces causal dependencies between agent actions and measures what fraction of a team's effort actually contributed to the outcome.

Abstract

Existing multi-agent benchmarks primarily test in competitive settings, short-horizon interactions under 20 steps, or simply aggregate individual performance, failing to isolate and highlight genuine collaboration capabilities of LLM-based agents. We introduce AgentWorld, a benchmark of 100 human-annotated tasks (with 100 augmented variants) for evaluating long-horizon, multi-agent collaboration. Tasks span 50+ interaction rounds across a rich MMORPG sandbox and require 3-20 agents with asymmetric roles and abilities to coordinate through communication, joint planning, and resource sharing under a blackbox setting where each agent acts independently without access to others'internal states. To quantify collaboration effectiveness in addition to conventional binary task success, we propose Causal Collaboration Effectiveness (CCE), a graph-based metric that traces causal dependencies between agent actions and measures what fraction of a team's effort actually contributed to the outcome. Experiments with Gemini 3 Flash, Claude Haiku 4.5, GPT-5 Mini, and DeepSeek R1-70B show that even the best model achieves only 52.0% task success, with systematic failure modes including communication breakdowns, role confusion, and inability to maintain shared plans across rounds. AgentWorld is fully open-source.

View source

Similar papers

#artificial intelligence Preprint Sep 2026

Rethinking Multi-Agent Collaboration: When More Is Less

SAIGE, a lightweight multi-agent collaboration mechanism based on Semantic-Aware Incremental Graph Evolution, is proposed, suggesting that multi-agent superiority is bounded by task structure rather than universal, and that more agents do not necessarily make a system more intelligent.

Yizhen Yuan, Yi-Bo Wu, Yi-Han Zhang et al. · 0 citations
#artificial intelligence Preprint Sep 2026

ParaAgent: Reinforcing Parallel Acting in Open-World Tool Environments

Language model agents are increasingly deployed in open-world tool environments, which require balancing exploring unknown capabilities and exploiting known ones. Existing methods face a performance-efficiency tradeoff: they either rigidly decouple exploration and execution or interleave them without coordination. We a...

Shengbin Yue, Hongru Wang, Siyuan Wang et al. · 0 citations
Preprint Aug 2026

CoCoBench: A Cooperative Coordination Benchmark for Embodied Multi-Agent Task Planning

CoBench is introduced, a construct-level benchmark for evaluating multi-agent embodied coordination in executable household tasks and shows that coordination ability is highly construct-specific: strong overall performance does not imply balanced competence across different coordination types.

Yang Chen, Ye-Xin Xie, Li-Rong Che et al. · 0 citations
#artificial intelligence Preprint Sep 2026

MASkills: Continual Skills Optimization for Multi-Agent LLM Systems

MASkills presents a new agent-optimization pipeline that integrates skill-conditioned credit assignment, hierarchical credit aggregation, and momentum-smoothed optimization, enabling agent skill libraries to evolve through refinement, induction, consolidation, and pruning.

Huaiyuan Yao, Xiaoou Liu, Charles Fleming et al. · 1 citation
#machine learning Preprint Sep 2026

Rethinking the Evaluation of Efficiency Methods for Multi-Agent Systems

This work introduces a controlled and MAS-demanding diagnostic benchmark for representative MAS efficiency methods and shows that many reported gains are setup-dependent and may arise from structural collapse, disabled tool pathways, or starting systems where random pruning already preserves accuracy, rather than robus...

Jia-Mu Zhang, Ling-Xi Zhang, Peng-Jun Lu et al. · 1 citation
#artificial intelligence Preprint Sep 2026

CollabFlow: Recursive Self-Improvement of Agent Collaboration

Recursive self-improvement (RSI) lets a system improve from its own outcomes; in LLM-based multi-agent systems, Agents refine one another within a task, and outcomes improve how they collaborate across tasks. However, existing multi-agent collaboration leaves this loop open: collaboration is pre-defined at the operator...

Xiao Huang, Ming-Da Zhang, Jun-Ming Zhang et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.