Multi agent PPO for terahertz cell free mobile edge computing networks
This work develops a multi-agent proximal policy optimization (MAPPO) algorithm to achieve a long-term optimal tradeoff between energy consumption and latency and demonstrates that under identical hyperparameter settings and training episodes, MAPPO achieves stable convergence and outperforms other MADRL baselines.