AsynCodeBench is introduced, a dependency-centric benchmark for asynchronous multi-agent software engineering that represents each task with an explicit dependency graph and executable Dependency Checkers and proposes two complementary measures: Asynchronous Dependency Pass Rate (ADPR), which measures how many cross-agent dependencies are ultimately satisfied, and Dependency Resolution Step (DRS), which measures when each dependency first becomes satisfied during execution.
Abstract
Multi-agent coding has emerged as an increasingly active direction in software engineering, where complex development tasks are decomposed across multiple specialized agents working on different parts of the problem. Despite the shift from individual problem solving to distributed collaboration, multi-agent systems still lack a direct measure of collaboration and are largely evaluated through task-level outcomes inherited from single-agent coding, conflating individual coding capability with cross-agent coordination. We introduce AsynCodeBench, a dependency-centric benchmark for asynchronous multi-agent software engineering that represents each task with an explicit dependency graph and executable Dependency Checkers. Through this dependency-tracking process, we propose two complementary measures: Asynchronous Dependency Pass Rate (ADPR), which measures how many cross-agent dependencies are ultimately satisfied, and Dependency Resolution Step (DRS), which measures when each dependency first becomes satisfied during execution. AsynCodeBench comprises 19 tasks from real-world repositories, exposing 52 directed dependencies as explicit units for evaluating cross-agent collaboration. Experiments across model families, scales, and generations reveal a clear gap between coding and collaboration capability: improvements in coding performance do not necessarily translate into stronger collaboration, and task-level metrics can diverge substantially from dependency-level collaboration measures. Dependency-trajectory analysis further reveals that successful coordination often emerges not gradually, but through concentrated bursts in which many dependencies become resolved over a short portion of the execution trajectory, a pattern we term a hopping window.
Large language models (LLMs) have progressively evolved into the core of autonomous agents. Building on this progress, LLM-based multi-agent systems (MAS) coordinate multiple agents into a synergistic team to accomplish complex tasks that exceed the capabilities of individual agents. The effectiveness of such systems d...
Qi-Zhi Chu, Ze-Kai Yu, Si-Jie Wen et al.· 0 citations
As AI agents increasingly tackle complex repository-level coding tasks, distributing work across multiple agents is a natural way to scale beyond the capabilities of a single agent. To coordinate their interdependent work, these agents share findings and agree on interfaces between modules. However, exchanged informati...
Jia-Qi Xue, Yan-Jun Wang, Xiangci Li et al.· 0 citations
Agents are increasingly considered first-class participants in the software development lifecycle. However, architectural choices that govern the collaboration of agents are fixed inside the frameworks that implement multi-agent systems, leading to difficulty of measuring the individual contribution of these architectu...
Konsta Kalliokoski, François Christophe, Henri Kärkkäinen et al.· Proceedings of the ACM/IEEE...· 0 citations
Multi-agent coding systems are designed to tackle complex software engineering tasks through collaboration. However, existing evaluations typically assume configured organizations are followed faithfully, whereas reality differs. This behavioral gap, combined with differences in underlying system components, prevents c...
This work synthesizes a task-conditioned temporal workflow graph that jointly specifies agent connectivity and edge-level communication semantics, and introduces ReActNet, a training-free framework that compiles a query and a set of role-specialized agents into a sequence of directed communication graphs.
Recently, the rapid development of large language models (LLMs) has reshaped software engineering by enabling autonomous code agents that plan, execute, and utilize external tools iteratively to tackle complex tasks. Beyond achieving functional correctness, these agents must faithfully follow process instructions and c...
Bo-Si Wen, Cun-Xiang Wang, Jia-Yi Gui et al.· 0 citations
With $2.1 million funding from Google.org, the open-source Public Transit Intelligence Hub will unify public transit monitoring, operations, and passenger communication.
Professor Sherry Turkle’s new book, “Artificial Intimacy,” offers a withering critique of chatbots and the antisocial dynamics she believes they encourage.
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.