Skip to content
Preprint

Planner-Conditioned Diffusion for Coordinated Multi-Agent Exploration

Aug 2026 · 0 citations · 21 references
Computer Science

TL;DR

A Planner-Conditioned Diffusion Policy (PCDP) is proposed, trained on demonstrations from multiple planner styles with planner identity as an explicit conditioning input, enabling a single shared model to learn a multimodal trajectory distribution and generate diverse, controllable trajectory candidates from the same observation.

Abstract

Coordinated multi-agent exploration requires not only efficient individual coverage but also non-redundant coverage across agents over extended planning horizons. Conventional approaches rely on hand-crafted coordination rules, while end-to-end multi-agent learning methods are difficult to scale and train. Diffusion-based planners such as DARE offer a promising alternative by generating long-horizon trajectories instead of single-step actions, but existing methods are trained on a narrow planner distribution, limiting behavioral diversity and inference-time controllability. We propose a Planner-Conditioned Diffusion Policy (PCDP) for graph-based multi-agent exploration. PCDP is trained on demonstrations from multiple planner styles with planner identity as an explicit conditioning input, enabling a single shared model to learn a multimodal trajectory distribution and generate diverse, controllable trajectory candidates from the same observation. Rather than learning coordination end-to-end, we reuse this multimodal single-agent policy across all agents and introduce coordination through local reranking, in which nearby agents jointly select the trajectory combination with minimal predicted overlap. We evaluate PCDP against classical and diffusion-based baselines on 100 held-out maps in a four-agent simulation setting. PCDP matches the perfect success rate of the diffusion-based baselines while improving mean max-agent travel, total team travel, and agent imbalance. Crucially, reranking alone over a single-planner baseline yields only marginal gains, indicating that planner-conditioned multimodality is the main contributor to improved coordination. Qualitative simulation results and real-robot experiments with two agents further validate that diverse long-horizon trajectory generation produces emergent spatial separation between agents without any explicit repulsion mechanism.

View source

Similar papers

Preprint Aug 2026

Search-Aided Joint Agent-Environment Reinforcement Learning for Robust Lifelong Multi-Agent Path Finding with Rotations

This work introduces a unified RL formulation that jointly optimizes agent and environment policies, where the environment policy learns graph edge costs to provide global movement guidance via backward Dijkstra search and achieves significant improvements over the strong search-based planner, Causal-PIBT, across multi...

He Jiang, Jingtian Yan, Yulun Zhang et al. · 0 citations
#artificial intelligence Open access Aug 2026

Generalizable Multi-Agent Planning From Signal Temporal Logic Specifications via Diffusion

A new diffusion method for multi-agent planning with STL specifications is introduced, making the approach generalizable to novel formulas whose predicates are placed anywhere within the goal region covered during training, while achieving the same scalability as existing learning-based methods.

Joe Eappen, Zikang Xiong, S. Iyengar et al. · 0 citations
Preprint Aug 2026

PlanPO: Group Planning-Aware Policy Optimization for Multi-Turn Agentic LLMs

Group Planning-aware Policy Optimization (PlanPO) is proposed, a simple yet effective RL method for learning generalizable planning abilities beyond task-specific high-quality behavior patterns that enables agents to actively learn generalizable and deliberate behaviors spanning interaction planning and textual generat...

D. Liang, Liyuan He, Xuan Feng et al. · 0 citations
Preprint Aug 2026

Discovering Diverse Planning Policies for Multimodal Embodied Agents with Quality-Diversity Optimization

Multimodal embodied agents are increasingly required to solve long-horizon tasks by integrating visual observations, textual goals, and interaction history into closed-loop decision making. However, state-of-the-art large-model-based planners often rely on a single dominant planning style during execution. Once this ex...

Peng Xu, Yong Liu, Xiaoya Nan et al. · 0 citations
Open access Aug 2026

Decentralized Model-Based ACKTR for Large-Scale Multi-Agent Path Planning Under Partial Observability

This work formulate large-scale MAPP as a partially observable networked Markov decision process as a decentralized model-based Actor-Critic using the Kronecker-factored trust region (DM-ACKTR) algorithm, which consistently obtains the highest TCR and lowest CR.

Ye-Min Liu, Jinhao Yang, Xiang-Yu Ma et al. · 0 citations
Preprint Aug 2026

Diffusion Policies for Short-Horizon Planning in Robot Crowd Navigation

Robot crowd navigation requires safe and efficient decision-making under dense, dynamic, and multimodal human--robot interactions. Existing reinforcement-learning methods typically output a single reactive action at each timestep, which limits their ability to represent diverse short-term avoidance strategies. We propo...

Wen-Dong Li, J. Garcke · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.