To quantify collaboration effectiveness in addition to conventional binary task success, Causal Collaboration Effectiveness (CCE) is proposed, a graph-based metric that traces causal dependencies between agent actions and measures what fraction of a team's effort actually contributed to the outcome.
Abstract
Existing multi-agent benchmarks primarily test in competitive settings, short-horizon interactions under 20 steps, or simply aggregate individual performance, failing to isolate and highlight genuine collaboration capabilities of LLM-based agents. We introduce AgentWorld, a benchmark of 100 human-annotated tasks (with 100 augmented variants) for evaluating long-horizon, multi-agent collaboration. Tasks span 50+ interaction rounds across a rich MMORPG sandbox and require 3-20 agents with asymmetric roles and abilities to coordinate through communication, joint planning, and resource sharing under a blackbox setting where each agent acts independently without access to others'internal states. To quantify collaboration effectiveness in addition to conventional binary task success, we propose Causal Collaboration Effectiveness (CCE), a graph-based metric that traces causal dependencies between agent actions and measures what fraction of a team's effort actually contributed to the outcome. Experiments with Gemini 3 Flash, Claude Haiku 4.5, GPT-5 Mini, and DeepSeek R1-70B show that even the best model achieves only 52.0% task success, with systematic failure modes including communication breakdowns, role confusion, and inability to maintain shared plans across rounds. AgentWorld is fully open-source.
SAIGE, a lightweight multi-agent collaboration mechanism based on Semantic-Aware Incremental Graph Evolution, is proposed, suggesting that multi-agent superiority is bounded by task structure rather than universal, and that more agents do not necessarily make a system more intelligent.
Yizhen Yuan, Yi-Bo Wu, Yi-Han Zhang et al.· 0 citations
Language model agents are increasingly deployed in open-world tool environments, which require balancing exploring unknown capabilities and exploiting known ones. Existing methods face a performance-efficiency tradeoff: they either rigidly decouple exploration and execution or interleave them without coordination. We a...
Shengbin Yue, Hongru Wang, Siyuan Wang et al.· 0 citations
CoBench is introduced, a construct-level benchmark for evaluating multi-agent embodied coordination in executable household tasks and shows that coordination ability is highly construct-specific: strong overall performance does not imply balanced competence across different coordination types.
Yang Chen, Ye-Xin Xie, Li-Rong Che et al.· 0 citations
MASkills presents a new agent-optimization pipeline that integrates skill-conditioned credit assignment, hierarchical credit aggregation, and momentum-smoothed optimization, enabling agent skill libraries to evolve through refinement, induction, consolidation, and pruning.
Huaiyuan Yao, Xiaoou Liu, Charles Fleming et al.· 1 citation
This work introduces a controlled and MAS-demanding diagnostic benchmark for representative MAS efficiency methods and shows that many reported gains are setup-dependent and may arise from structural collapse, disabled tool pathways, or starting systems where random pruning already preserves accuracy, rather than robus...
Jia-Mu Zhang, Ling-Xi Zhang, Peng-Jun Lu et al.· 1 citation
Recursive self-improvement (RSI) lets a system improve from its own outcomes; in LLM-based multi-agent systems, Agents refine one another within a task, and outcomes improve how they collaborate across tasks. However, existing multi-agent collaboration leaves this loop open: collaboration is pre-defined at the operator...
Xiao Huang, Ming-Da Zhang, Jun-Ming Zhang et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.