Skip to content

OpenMAS-GCom. A Diagnostic Benchmark for Graph-enhanced Multi-Agent Systems

Sep 2026 · 0 citations · 41 references
Computer Science

TL;DR

OpenMAS-GCom is introduced, a benchmark for diagnosing how these components affect G-MAS performance through controlled interventions, and 400 G-MAS-Complex tasks requiring agents to combine information from multiple documents, resolve conflicting records, and return specified values with source identifiers are added.

Abstract

Graph-enhanced multi-agent systems (G-MAS) coordinate large language model agents through communication graphs and role assignments, which determine how agents exchange information and divide responsibilities. However, final-score comparisons across systems combine differences in models, communication patterns, roles, and computation costs, making performance differences difficult to attribute to specific communication structures, role assignments, and information flows. To address this evaluation attribution problem, we introduce OpenMAS-GCom, a benchmark for diagnosing how these components affect G-MAS performance through controlled interventions. We represent systems through collaboration units, communication links, shared intermediate information, and execution rules. OpenMAS-GCom compares original systems with versions modified by changing one component while keeping tasks, models, prompts, and budget limits fixed. We rewire communication edges, remove specialist or critic agents, replace intermediate messages with incorrect content, and disable workers during execution. The benchmark evaluates 17 single-agent, ordinary multi-agent, and graph-enhanced configurations on 29 datasets across six domains. We add 400 G-MAS-Complex tasks requiring agents to combine information from multiple documents, resolve conflicting records, and return specified values with source identifiers. Experiments show larger mean losses after specialist removal than after critic removal, different performance degradation under incorrect messages and worker failures despite similar original scores, and different configurations achieving the highest accuracy and accuracy per token on G-MAS-Complex.

View source

Similar papers

Preprint Aug 2026

ForestBench: A Unified Graph Framework for Evaluating Multi-Agent Collaboration

This work introduces a generalizable evaluation framework that maps native MAS traces into a shared space of unified collaboration graphs, enabling different methods to be evaluated under the same representation, reference set, and metric panel.

Guo Chen, Ziwen Li, Reed Li et al. · 0 citations
#machine learning Preprint Sep 2026

GraphMAS: A Systematic Benchmark of Multi-Agent Coordination for Graph Learning

LLM-based multi-agent systems coordinate specialized reasoning through aggregation, interaction, and adaptive control, yet their potential for graph learning remains unexplored. Graph learning is a natural setting for such systems because useful evidence may arise from heterogeneous local, long-range, global structural...

Jia-Yi Yang, Yi-Fang Chen, Yuan-Fu Sun et al. · 0 citations
#machine learning Preprint Sep 2026

MACE: Memory-Agent Co-Evolution with Adaptive Memory Graphs for Multi-Agent Systems

LLM-based multi-agent systems generate collaboration traces that record how agents plan tasks, verify intermediate results, and repair failures. Reusing these procedures requires preserving an action's prerequisites and the outputs needed by subsequent agents. Our empirical studies show that grouping these dependencies...

Kai-Rui Yang, Ming-Hao An, Xun-Kai Li et al. · 0 citations
#artificial intelligence Preprint Sep 2026

Inference-Time Graph Engineering for Multi-Agent LLM Workflows

Recent multi-agent LLM systems increasingly rely on graph-structured communication to coordinate specialized agents. We revisit multi-agent orchestration from a graph-engineering perspective: rather than optimizing a static topology, we synthesize a task-conditioned temporal workflow graph that jointly specifies agent...

Katherine Tieu, Dong-Qi Fu, Ying-Long Xia et al. · 1 citation · ⚡1
Preprint Aug 2026

Discovering Efficient and Explainable Communication Topologies for LLM-based Multi-Agent Systems via Causal Inference

E2-Explainer is proposed, a model-agnostic framework for providing interpretable explanations of communication topologies produced by arbitrary topology generators that identifies compact communication subgraphs supported by edge-level evidence of task preservation.

Jun-Zhi Li, Peng He, Qirui Ji et al. · 1 citation
#machine learning Preprint Sep 2026

Rethinking the Evaluation of Efficiency Methods for Multi-Agent Systems

This work introduces a controlled and MAS-demanding diagnostic benchmark for representative MAS efficiency methods and shows that many reported gains are setup-dependent and may arise from structural collapse, disabled tool pathways, or starting systems where random pruning already preserves accuracy, rather than robus...

Jia-Mu Zhang, Ling-Xi Zhang, Peng-Jun Lu et al. · 1 citation

Related blog posts

Microsoft Research Blog Sep 30, 2026

Forecasting space weather risks on power grids

Extreme space-weather events can damage power systems on Earth and degrade GPS accuracy and satellite operations. A new machine learning system can predict where damage is likely to occur 30-60 minutes before a storm arrives. The post Forecasting space weather risks on power grids appeared first on Microsoft Research.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.