Skip to content
Conference

SPD-MAPPO: Reinforcement Learning With Stochastic Policy Distillation for Multi-Vehicle Coordination in Open-pit Mine

Jul 2026 · 2026 IEEE/ASME International Conference on Advanced Intelligent Mechatronics (AIM) · pp. 1-6 · 0 citations · 17 references

Abstract

Coordination of multiple autonomous trucks is crucial for enhancing the efficiency and safety of modern mining, yet it is challenged by dynamic vehicle-to-vehicle interactions and the complexity of mining transportation. Conventional rule-based methods struggle to balance efficiency with success rates and lack flexibility in diverse scenarios. While multi-agent reinforcement learning (MARL) shows great promise for cooperative tasks, its application in real world is often hampered by challenges in convergence. To address these challenges, we propose SPD-MAPPO, a novel multi-stage learning framework that transfers expertise from imitation learning(IL) to cooperative ability in MARL. The framework first employs IL to pre-train a policy with basic single-agent driving ability, which is subsequently refined for cooperative behaviors through MARL. Specifically, we design a Stochastic Policy Distillation (SPD) mechanism to bridge the gap between single-agent expertise and multi-agent coordination, and a multi-head critic network to achieve more precise credit assignment. We validate our method in a high-fidelity simulator with a real-world map of mine and a truck dynamics model. Our method outperforms typical rule-based and MARL methods in success rate, efficiency, and operational accuracy.

View source

Similar papers

Review Open access Aug 2026

Advances in Multi-Agent Deep Reinforcement Learning: Methods with Applications and Challenges

This paper presents a narrative survey of recent developments in MARL and examines research directions centred on centralised training with decentralised execution (CTDE), value decomposition, learned communication, graph-based methods, and model-based learning.

Abdur Rakib, K. Phung, Marco Pérez Hernández et al. · 0 citations
Open access Aug 2026

AGTA: Topology-Aware Sequential Decision-Making in Multi-Agent Reinforcement Learning

Action Generation with Topology Awareness (AGTA), a topology-aware sequential decision-making framework in MARL that integrates inter-agent correlation modeling with topology-guided decision-order optimization, and outperforms the state-of-the-art counterparts.

Kun Hu, Shanghua Wen, Wendi Wu et al. · 0 citations
2025

Learning and Planning Multi-Agent Tasks via an MoE-based World Model

M3W is a novel approach that applies mixture-of-experts (MoE) to world model instead of policy, enabling both learning and planning, and demonstrates superior performance, sample efficiency, and multi-task adaptability.

Zi-Jie Zhao, Zhong Zhao, Kaixuan Xu et al. · 9 citations
Jul 2026

Efficient Heterogeneous Exploration with Mutual Policy Divergence Maximization for Multiagent Reinforcement Learning.

This work introduces a novel MARL framework, Multi-Agent Divergence Policy Optimization (MADPO) with Mutual Policy Divergence Maximization (Mutual PDM), and proposes a new extension of CCS divergence for measuring policy divergence of more than two agents, the Generalized Conditional Cauchy-Schwarz (GCCS) divergence.

Haowen Dou, Lujuan Dang, Mingfei Lu et al. · 0 citations
Aug 2026

CA-MARL: credit-aware multi-agent reinforcement learning for USV pursuit–evasion missions under complex maritime disturbances

Multi-unmanned surface vehicle (USV) pursuit–evasion missions in maritime environments presents significant challenges due to dynamic ship populations, high-dimensional observations, and the gap between idealised simulations and real-world maritime physics. To address these challenges, we propose a Credit-Aware Multi-Agent Reinforcement Learning (CA-MARL) framework for multi-USV pursuit–evasion. The framework features two key innovations: a Residual Self-Attention module that adapts to varying fleet sizes through permutation-invariant attention, and a Mixed Credit Assignment module that enhances centralised value estimation with decentralised branches. Moreover, to bridge the simulation-to-reality gap, we develop a high-fidelity 3D virtual platform using Unity3D that incorporates maritime factors, such as hydrodynamics and wave disturbances, which are typically overlooked in USV simulations but critical for maritime operations. Experiments demonstrate that our method achieves superior coordination, sample efficiency, and policy robustness compared to existing baselines, providing a credible foundation for deploying MARL policies in realistic multi-USV scenarios.

Shunyu Tian, Weiyu Tao, Xiangyu Wu et al. · 0 citations

Towards Streamlined Learning and Search for Multi-Agent Optimization

Focusing on multi-agent path finding as an exemplary problem, this paper proposes to simplify two popular approaches to MAPF, namely multi-agent reinforcement learning and adaptive search, to enable seamless combination and transferability of methods without substantial engineering effort.

Thomy Phan · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.