Multi-Agent Reinforcement Learning for Movable Antenna-aided Cell-Free Massive MIMO Systems
This work proposes the graph-based learning individual intrinsic reward heterogeneous-agent proximal policy optimization (GLIIR-HAPPO) algorithm, a novel heterogeneous multi-agent reinforcement learning (MARL) framework that fundamentally overcomes this impasse by systematically decomposing the original coupled optimiz...