Skip to content

AsynCodeBench: Benchmarking Collaboration of Asynchronous Multi-Agent Systems in Software Engineering

Sep 2026 · 0 citations · 45 references
Computer Science

TL;DR

AsynCodeBench is introduced, a dependency-centric benchmark for asynchronous multi-agent software engineering that represents each task with an explicit dependency graph and executable Dependency Checkers and proposes two complementary measures: Asynchronous Dependency Pass Rate (ADPR), which measures how many cross-agent dependencies are ultimately satisfied, and Dependency Resolution Step (DRS), which measures when each dependency first becomes satisfied during execution.

Abstract

Multi-agent coding has emerged as an increasingly active direction in software engineering, where complex development tasks are decomposed across multiple specialized agents working on different parts of the problem. Despite the shift from individual problem solving to distributed collaboration, multi-agent systems still lack a direct measure of collaboration and are largely evaluated through task-level outcomes inherited from single-agent coding, conflating individual coding capability with cross-agent coordination. We introduce AsynCodeBench, a dependency-centric benchmark for asynchronous multi-agent software engineering that represents each task with an explicit dependency graph and executable Dependency Checkers. Through this dependency-tracking process, we propose two complementary measures: Asynchronous Dependency Pass Rate (ADPR), which measures how many cross-agent dependencies are ultimately satisfied, and Dependency Resolution Step (DRS), which measures when each dependency first becomes satisfied during execution. AsynCodeBench comprises 19 tasks from real-world repositories, exposing 52 directed dependencies as explicit units for evaluating cross-agent collaboration. Experiments across model families, scales, and generations reveal a clear gap between coding and collaboration capability: improvements in coding performance do not necessarily translate into stronger collaboration, and task-level metrics can diverge substantially from dependency-level collaboration measures. Dependency-trajectory analysis further reveals that successful coordination often emerges not gradually, but through concentrated bursts in which many dependencies become resolved over a short portion of the execution trajectory, a pattern we term a hopping window.

View source

Similar papers

#artificial intelligence Preprint Oct 2026

MASBench: Benchmarking LLM-based Multi-Agent Collaboration under Partial Observability

Large language models (LLMs) have progressively evolved into the core of autonomous agents. Building on this progress, LLM-based multi-agent systems (MAS) coordinate multiple agents into a synergistic team to accomplish complex tasks that exceed the capabilities of individual agents. The effectiveness of such systems d...

Qi-Zhi Chu, Ze-Kai Yu, Si-Jie Wen et al. · 0 citations
#natural language process... Preprint Oct 2026

AECP: Artifact-Exclusive Communication Protocol for Multi-Agent Code Generation

As AI agents increasingly tackle complex repository-level coding tasks, distributing work across multiple agents is a natural way to scale beyond the capabilities of a single agent. To coordinate their interdependent work, these agents share findings and agree on interfaces between modules. However, exchanged informati...

Jia-Qi Xue, Yan-Jun Wang, Xiangci Li et al. · 0 citations
Book Open access Oct 2026

KERKIS: a Modeling Language for Multi-Agent System Interactions in AI-Native Software Engineering

Agents are increasingly considered first-class participants in the software development lifecycle. However, architectural choices that govern the collaboration of agents are fixed inside the frameworks that implement multi-agent systems, leading to difficulty of measuring the individual contribution of these architectu...

Konsta Kalliokoski, François Christophe, Henri Kärkkäinen et al. · 0 citations
#artificial intelligence Preprint Sep 2026

OpenCollab: A Multi-Agent Coding Framework with Programmable Collaboration and Controllable Runtime

Multi-agent coding systems are designed to tackle complex software engineering tasks through collaboration. However, existing evaluations typically assume configured organizations are followed faithfully, whereas reality differs. This behavioral gap, combined with differences in underlying system components, prevents c...

Chun-Wah Hsu, Kai Gong, Yu Wu et al. · 0 citations
#artificial intelligence Preprint Sep 2026

Inference-Time Graph Engineering for Multi-Agent LLM Workflows

This work synthesizes a task-conditioned temporal workflow graph that jointly specifies agent connectivity and edge-level communication semantics, and introduces ReActNet, a training-free framework that compiles a query and a set of role-specialized agents into a sequence of directed communication graphs.

Katherine Tieu, Dong-Qi Fu, Ying-Long Xia et al. · 1 citation · ⚡1
#natural language process... Preprint Sep 2026

MTAC-IFBench: Benchmarking Instruction-Following in Multi-Turn Agentic Coding

Recently, the rapid development of large language models (LLMs) has reshaped software engineering by enabling autonomous code agents that plan, execute, and utilize external tools iteratively to tackle complex tasks. Beyond achieving functional correctness, these agents must faithfully follow process instructions and c...

Bo-Si Wen, Cun-Xiang Wang, Jia-Yi Gui et al. · 0 citations

Related blog posts

MIT News · Artificial Intelligence Sep 29, 2026

Who we become when we talk to machines

Professor Sherry Turkle’s new book, “Artificial Intimacy,” offers a withering critique of chatbots and the antisocial dynamics she believes they encourage.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.