Multi-State Parameter Self-Optimization Using Deep Reinforcement Learning for Energy-Efficient Low-Latency Massive MIMO Systems
Abstract
Massive MIMO systems require simultaneous optimization of energy efficiency, latency, and handover performance, yet existing approaches address these objectives in isolation across disparate parameter spaces. This paper proposes a multi-state parameter self-optimization framework that jointly optimizes across five interdependent operational states—channel, mobility, system configuration, power, and latency—using deep reinforcement learning. We formulate the problem as a multi-objective Markov decision process and implement five optimization approaches: Hybrid Action Space Reinforcement Learning, Q-Learning with Kalman Filter prediction, LSTM Autoencoder for PAPR reduction, bio-inspired Integrated Fruit Fly Salp Swarm Optimization for power allocation, and a proposed Multi-Agent Deep Q-Network (MA-DQN) with experience replay. Simulation results across antenna configurations from 16 to 256 elements and user counts from 5 to 40 show that the proposed MA-DQN achieves a composite performance score of $83 \pm 1.8 / 100$ across all five states (averaged over 10 seeded runs), outperforming the best single-objective method by $\mathbf{2 6} \boldsymbol{\%}$. The framework delivers 29-73% energy efficiency improvement over fixed baselines, with the learned policy favoring moderate power (0.1-0.5W) and lower antenna counts (16-32)—consistent with analytical models that show circuit power dominance at high antenna counts.