2026· IEEE Transactions on Wireless Communications· Vol 25, pp. 21403-21417· 0 citations· 46 references
Computer Science
Abstract
Multi-agent deep reinforcement learning (DRL) offers a promising framework for inter-cell interference mitigation in multi-cell networks. In such networks, each cell is associated with an agent that learns from its local environment to maximize a reward, such as spectral efficiency. To effectively mitigate inter-cell interference, agents typically share model weights or local experiences with one another or with a central node. However, the exchange of such information incurs significant communication overhead in each communication round between the central node and the individual agents, posing a major bottleneck to efficient multi-agent DRL-based inter-cell interference mitigation. This paper presents a novel dimension-independent multi-agent DRL algorithm for multi-cell interference mitigation. By leveraging zeroth-order optimization, the proposed algorithm reduces the communication overhead from <inline-formula> <tex-math notation="LaTeX">$\mathcal {O}(d)$ </tex-math></inline-formula> to <inline-formula> <tex-math notation="LaTeX">$\mathcal {O}(1)$ </tex-math></inline-formula>, where <inline-formula> <tex-math notation="LaTeX">$d$ </tex-math></inline-formula> denotes the shared information dimension. This is achieved by exchanging only a constant number of scalar values between the central node and the agents in each communication round, independent of the dimension <inline-formula> <tex-math notation="LaTeX">$d$ </tex-math></inline-formula> of the shared weights or experiences. The proposed algorithm is evaluated on millimeter-wave networks with varying numbers of cells, demonstrating its effectiveness for interference mitigation. Specifically, under universal frequency reuse, the total sum-rate increases almost linearly with the number of cells. Simulation results show that the proposed algorithm effectively mitigates interference and maximizes spectral efficiency in line-of-sight (LoS), non-LoS, and mixed environments, while significantly reducing communication overhead.
Numerical evaluations reveal that the multi-agent DDPG approach substantially outperforms single-agent in dense scenarios, and the multi-agent demonstrates highly efficient computational convergence of dense scenarios with $400$ users.
Omar Rady, Mohamed Ayman, Ali Arafa et al.· 0 citations
Artificial intelligence (AI) and machine learning (ML) are increasingly applied to wireless and cellular networks. With sixth-generation (6G) systems envisioned as AI-native, reinforcement learning (RL) offers a natural approach to complex network management and operation. This paper focuses on user admission control in multi-cell massive multiple-input multiple-output (MIMO) systems, where naive selfish strategies aiming to maximize local sum-rate can trigger a tragedy of the commons, degrading per-user performance and generating severe inter-cell interference (ICI). To address these challenges, we introduce a structured RL framework for massive MIMO systems. In particular, the policy is structured to introduce physical inductive bias terms, such as an interference-sensitive attenuation factor, which enables interference-aware learning through the open radio access network (O-RAN) architecture. Through stability analysis, we show that such physical inductive bias terms can guarantee network-wide stability. Experimental results demonstrate that the proposed approach balances aggregate spectral efficiency with per-user performance and maintains robustness during traffic surges, whereas selfish strategies suffer from degraded per-user performance.
Jinho Choi· IEEE Transactions on Communi...· 0 citations
Multi-Agent Systems (MAS) have emerged as a powerful paradigm for modeling complex interactions among autonomous entities in distributed environments. In Multi-Agent Reinforcement Learning (MARL), communication enables coordination but can lead to inefficient information exchange, since agents may generate redundant or non-essential messages. While prior work has focused on boosting task performance with information exchange, the existing research lacks a thorough investigation of both the appropriate definition and the optimization of communication protocols (communication topology and message). To fill this gap, we introduce a unified framework for learning multi-round communication protocols that are both effective and efficient. Within this framework, we propose three novel Communication Efficiency Metrics (CEMs) to guide and evaluate the learning process: the Information Entropy Efficiency Index (IEI) and Specialization Efficiency Index (SEI) for efficiency-augmented optimization, and the Topology Efficiency Index (TEI) for explicit evaluation. We integrate IEI and SEI as the adjusted loss functions to promote informative messaging and role specialization, while using TEI to quantify the trade-off between communication volume and task performance. Through comprehensive experiments, we demonstrate that our learned communication protocols can significantly enhance communication efficiency and achieves better cooperation performance with improved success rates.
Action Generation with Topology Awareness (AGTA), a topology-aware sequential decision-making framework in MARL that integrates inter-agent correlation modeling with topology-guided decision-order optimization, and outperforms the state-of-the-art counterparts.
Kun Hu, Shanghua Wen, Wendi Wu et al.· Mathematics· 0 citations
Sixth-generation (6G) mobile communication poses unprecedented challenges for resource scheduling under personalized demands. Cell-free massive multiple-input multiple-output (CF-mMIMO), with its user-centric characteristics, has emerged as a key technology for satisfying personalized demands. However, faced with heterogeneous quality-of-service (QoS) requirements, existing reinforcement learning schemes are constrained by partial observability, making it difficult to balance overall system performance and personalized demands. Consequently, we propose a graph-embedded multi-agent deep deterministic policy gradient (G-MADDPG) scheme. Guided by personalized demands, proposed G-MADDPG formulates a maximization problem for system weighted sum spectral efficiency and introduces differentiated QoS penalties. In addition, graph neural networks (GNNs) are embedded into the policy learning and value estimation processes of reinforcement learning, endowing agents with enhanced structural reception and cooperative capabilities. Simulation results demonstrate that proposed G-MADDPG scheme outperforms existing benchmark schemes in both convergence speed and performance evaluation.
Yu-Heng An· 2026 8th International Confe...· 0 citations
The integration of unmanned aerial vehicles (UAVs) as aerial base stations has emerged as a key enabler for next-generation wireless networks, particularly in disaster recovery, temporary events, and infrastructure-deficient regions. However, multi-UAV deployments introduce severe co-channel interference due to spectrum reuse and overlapping coverage areas, while existing spectrum allocation methods either rely on centralized optimization with limited scalability or on reinforcement learning frameworks that lack spatial awareness of interference sources. To address these challenges, this paper proposes a Hybrid DeepMUSIC-assisted Cooperative Multi-Agent Deep Reinforcement Learning (MADRL) framework for intelligent spectrum allocation and interference management in multi-UAV 6G networks. The proposed framework integrates a hybrid interference localization module, which fuses the classical MUltiple SIgnal Classification (MUSIC) algorithm with a deep neural network to accurately estimate the direction of arrival (DoA) of interference sources, into a DeepMUSIC-enhanced state representation used by cooperative Deep Q-Network (DQN) agents trained under a Centralized Training and Decentralized Execution (CTDE) paradigm, enabling coordinated yet fully distributed spectrum allocation decisions. Extensive simulations demonstrate that the proposed Hybrid DeepMUSIC module reduces the mean DoA estimation error to approximately 0.105°, more than an order of magnitude better than classical MUSIC and standalone DeepMUSIC estimators. Compared with seven baseline algorithms spanning heuristic, optimization-based, single-agent, and cooperative multi-agent reinforcement learning approaches, the proposed framework achieves the highest network throughput, SINR, spectrum efficiency, and energy efficiency, together with the fastest and most stable training convergence, reaching a stable cooperative reward of 76.246 within approximately 371 training epochs. The framework further maintains near-linear computational scaling with the number of UAV agents, confirming its suitability for real-time deployment in dense, AI-native multi-UAV 6G wireless communication systems.
Unknown authors· Technologies· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.