Traffic-Adaptive Per-Hop Multipath Routing in Multi-Hop UAV Networks
This work develops a multi-agent reinforcement learning (MARL) algorithm, termed Multi-Agent Proximal Policy Optimization with Dirichlet Modeling (MAPPO-DM), which follows the centralized-training-and-decentralized-execution framework and models continuous traffic-splitting actions using a Dirichlet distribution.
Zhen-Yu Zhao, Tiankui Zhang, Xiao-Xia Xu et al.
· 0 citations