2026· IEEE Transactions on Communications· Vol 74, pp. 12852-12865· 0 citations· 35 references
Computer Science
Abstract
Artificial intelligence (AI) and machine learning (ML) are increasingly applied to wireless and cellular networks. With sixth-generation (6G) systems envisioned as AI-native, reinforcement learning (RL) offers a natural approach to complex network management and operation. This paper focuses on user admission control in multi-cell massive multiple-input multiple-output (MIMO) systems, where naive selfish strategies aiming to maximize local sum-rate can trigger a tragedy of the commons, degrading per-user performance and generating severe inter-cell interference (ICI). To address these challenges, we introduce a structured RL framework for massive MIMO systems. In particular, the policy is structured to introduce physical inductive bias terms, such as an interference-sensitive attenuation factor, which enables interference-aware learning through the open radio access network (O-RAN) architecture. Through stability analysis, we show that such physical inductive bias terms can guarantee network-wide stability. Experimental results demonstrate that the proposed approach balances aggregate spectral efficiency with per-user performance and maintains robustness during traffic surges, whereas selfish strategies suffer from degraded per-user performance.
Sixth-generation (6G) mobile communication poses unprecedented challenges for resource scheduling under personalized demands. Cell-free massive multiple-input multiple-output (CF-mMIMO), with its user-centric characteristics, has emerged as a key technology for satisfying personalized demands. However, faced with heterogeneous quality-of-service (QoS) requirements, existing reinforcement learning schemes are constrained by partial observability, making it difficult to balance overall system performance and personalized demands. Consequently, we propose a graph-embedded multi-agent deep deterministic policy gradient (G-MADDPG) scheme. Guided by personalized demands, proposed G-MADDPG formulates a maximization problem for system weighted sum spectral efficiency and introduces differentiated QoS penalties. In addition, graph neural networks (GNNs) are embedded into the policy learning and value estimation processes of reinforcement learning, endowing agents with enhanced structural reception and cooperative capabilities. Simulation results demonstrate that proposed G-MADDPG scheme outperforms existing benchmark schemes in both convergence speed and performance evaluation.
Yu-Heng An· 2026 8th International Confe...· 0 citations
The proposed framework does not optimize only computational speed, but also clarifies the trade-off among execution time, SINR, spectral efficiency, and fairness under dynamic uplink CF-mMIMO conditions, indicating that this architecture serves as an adaptable platform to evaluate dynamic uplink power distribution across CF-mMIMO networks.
Hussein A. Jasim, M. F. A. Rasid, F. Hashim et al.· Engineer· 0 citations
: Massive multiple-input multiple-output (MIMO) technology is a key enabler for 5G and beyond wireless networks, offering significant improvements in spectral efficiency and link reliability. However, conventional beamforming techniques such as Zero Forcing (ZF) and Minimum Mean Square Error (MMSE) require complex matrix computations and fail to adapt efficiently to dynamic channel variations. To address these challenges, this paper proposes a Deep Deterministic Policy Gradient (DDPG)-based beamforming framework that formulates beamforming optimization as a continuous-action deep reinforcement learning problem. The proposed model directly generates complex-valued beamforming weight vectors to jointly maximize spectral efficiency (SE) and energy efficiency (EE) while minimizing the bit error rate (BER) and decision latency. An adaptive state representation incorporating channel state information (CSI), previous beamforming vectors, and performance metrics enables real-time policy learning under time-varying channel conditions. Simulation results demonstrate that the proposed method outperforms conventional and heuristic beamforming schemes, achieving up to 25% higher SE, 45% lower BER, 20% improvement in EE, and 30% latency reduction. The results validate the effectiveness of the proposed framework for energy-efficient, low-latency beamforming in next-generation massive MIMO wireless networks.
Nilakshee Rajule, Mithra Venkatesan, Harshada Magar et al.· Proceedings of the 1st Inter...· 0 citations
The Millimeter-wave massive Multiple-Input Multiple-Output (MIMO) is a core mechanism for the Sixth-Generation (6G) wireless communication networks. By including numerous antennas in the compact model of advanced smartphones, the MIMO enhances the network capacity and the Spectral Efficiency (SE). The growth of 6G technology is significant for the future and provided evolutionary and revolutionary solutions. The resource allocation in the MIMO-based wireless networks is selected for various users, aiming to optimize the network resource distribution. But the high increase in the antennas and users poses complexities for the resource allocation and interference suppression for the MIMO systems. In this research, an advanced Deep Reinforcement Learning (DRL)-based approach is proposed for efficient dynamic spectrum allocation in 6G MIMO systems. To perform spectrum allocation in 6G MIMO systems, an Adaptive Deep Multiagent Reinforcement Learning with Co-ordinate Attention (ADMRL-CA) model is developed. The DMRL mechanism is capable of handling varying traffic and channel conditions. The incorporation of the CA mechanism enhances the policy learning process for accurate spectrum allocation. The parameters of the ADMRL-CA are fine-tuned using the Flying workers phase Modified Termite Queen Algorithm (FMTQA). Finally, the model performance is analyzed with various existing models. The SE of the recommended FMTQA-ADMRL-CA is increased by 4.44% of DRL, 6.66% of SAC, 2.77% of DDPG and 10.88% of DMRL-CA when system uses as 32
nd
batch size. Hence, it is guaranteed that the recommended FMTQA-ADMRL-CA can allocate the spectrum efficiently and robustly in 6G MIMO systems than the existing methods.
Asha Aiyappan, Jafar A. Alzubi, M. P. Rajakumar et al.· Scientific Reports· 0 citations
A Deep Reinforcement Learning (DRL)-based framework for dynamic spectrum access in 6G heterogeneous Cognitive Radio Networks (Het-CRNs), wherein secondary users learn optimal channel selection policies through direct interaction with the radio environment, without requiring explicit statistical channel models is proposed.
Naadir Kamal, R. Kumar· Global Journal of Engineerin...· 0 citations
A reinforcement learning (RL)-driven control framework that operates over a bank of pretrained multi-rate AEs, each corresponding to a distinct compression ratio (CR), aiming to dynamically optimize the trade-off between reconstruction fidelity and signaling overhead.
Maryam Ansarifard, M. Sharma, Georgios Exarchakos et al.· arXiv.org· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.