Skip to content
Preprint

Certifying cooperation: a novel approach to cooperative multi-agent task generation

Sep 2026 · 0 citations
Computer Science

TL;DR

This framework exposes the gap between rewarded partial completion and realized cooperation by certifying what cooperation successful completion requires and using temporal cooperation graphs to reveal what policies exhibit.

Abstract

A shared reward gives agents a common objective, but leaves open when, how and even whether they must cooperate to succeed. We address these questions in the Laser Learning Environment, a multi-agent path-finding environment where cooperation materializes as one agent blocking a laser to let a teammate pass safely. We represent these interactions through temporal cooperation graphs whose timed edges connect helpers to beneficiaries, define six cooperation profiles as overlapping graph predicates, and prove that every cooperative trajectory satisfies at least one. By encoding the environment dynamics and profile predicates as propositional formulae, we distinguish tasks that admit}a profile in some winning trajectory from those that require it in every winning trajectory within a specified horizon. Used as filters, these queries turn a random layout sampler into a generator of tasks with certified cooperation requirements. Experiments with five multi-agent reinforcement learning algorithms show that training diversity improves joint success on unseen tasks when cooperation-free solutions exist. When cooperation is required, greater diversity improves individual-agent exits, but joint success remains near zero. Across five profile-certified pools, final exit rates averaged over algorithms separate the pools into four statistically distinguishable levels but this ordering primarily reflects partial completion: policies collect rewards for individual exits but rarely exhibit the profile required for joint success. Our framework exposes this gap between rewarded partial completion and realized cooperation by certifying what cooperation successful completion requires and using temporal cooperation graphs to reveal what policies exhibit.

View source

Similar papers

#artificial intelligence Preprint Sep 2026

Multi-Agent System Search via Active Substructure-aware Policy Optimization

LLMs enable multi-agent systems (MAS) to tackle complex tasks, but manually designing agent roles, prompts, and communication structures requires substantial expertise and effort. This motivates learning policies that construct query-specific MAS from execution reward. Existing approaches typically train these policies...

Bei-Cheng Xu, Bo-Wen Fan, Wei Qian et al. · 0 citations
#artificial intelligence Preprint Sep 2026

CollabFlow: Recursive Self-Improvement of Agent Collaboration

Recursive self-improvement (RSI) lets a system improve from its own outcomes; in LLM-based multi-agent systems, Agents refine one another within a task, and outcomes improve how they collaborate across tasks. However, existing multi-agent collaboration leaves this loop open: collaboration is pre-defined at the operator...

Xiao Huang, Ming-Da Zhang, Jun-Ming Zhang et al. · 0 citations
#natural language process... Preprint Sep 2026

MAS-OPD: On-Policy Distillation for Multi-agent Systems

MAS-OPD is presented, where Role-Advantage Specialization defines the role advantage as the difference between the teacher signals under target and non-target role conditions, and Privileged Attribution for Coordination attributes an interaction conflict to its source and supplies it to the teacher alone as privileged...

Qi-Yong Zhong, Mao Zheng, Ming-Yang Song et al. · 0 citations
#artificial intelligence Preprint Sep 2026

ParaAgent: Reinforcing Parallel Acting in Open-World Tool Environments

Language model agents are increasingly deployed in open-world tool environments, which require balancing exploring unknown capabilities and exploiting known ones. Existing methods face a performance-efficiency tradeoff: they either rigidly decouple exploration and execution or interleave them without coordination. We a...

Shengbin Yue, Hongru Wang, Siyuan Wang et al. · 0 citations
Preprint Sep 2026

AgentWorld: Benchmarking Long-Horizon Collaboration of Multi-agent LLMs

To quantify collaboration effectiveness in addition to conventional binary task success, Causal Collaboration Effectiveness (CCE) is proposed, a graph-based metric that traces causal dependencies between agent actions and measures what fraction of a team's effort actually contributed to the outcome.

Raphael Shu, Yu-Sen Zhang, Y. Cho et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.