Skip to content
Open access

Hybrid DeepMUSIC-Assisted Cooperative Multi-Agent Deep Reinforcement Learning for Intelligent Spectrum Allocation and Interference Management in Multi-UAV 6G Networks

Unknown authors
Sep 2026 · Technologies · 0 citations · 41 references

Abstract

The integration of unmanned aerial vehicles (UAVs) as aerial base stations has emerged as a key enabler for next-generation wireless networks, particularly in disaster recovery, temporary events, and infrastructure-deficient regions. However, multi-UAV deployments introduce severe co-channel interference due to spectrum reuse and overlapping coverage areas, while existing spectrum allocation methods either rely on centralized optimization with limited scalability or on reinforcement learning frameworks that lack spatial awareness of interference sources. To address these challenges, this paper proposes a Hybrid DeepMUSIC-assisted Cooperative Multi-Agent Deep Reinforcement Learning (MADRL) framework for intelligent spectrum allocation and interference management in multi-UAV 6G networks. The proposed framework integrates a hybrid interference localization module, which fuses the classical MUltiple SIgnal Classification (MUSIC) algorithm with a deep neural network to accurately estimate the direction of arrival (DoA) of interference sources, into a DeepMUSIC-enhanced state representation used by cooperative Deep Q-Network (DQN) agents trained under a Centralized Training and Decentralized Execution (CTDE) paradigm, enabling coordinated yet fully distributed spectrum allocation decisions. Extensive simulations demonstrate that the proposed Hybrid DeepMUSIC module reduces the mean DoA estimation error to approximately 0.105°, more than an order of magnitude better than classical MUSIC and standalone DeepMUSIC estimators. Compared with seven baseline algorithms spanning heuristic, optimization-based, single-agent, and cooperative multi-agent reinforcement learning approaches, the proposed framework achieves the highest network throughput, SINR, spectrum efficiency, and energy efficiency, together with the fastest and most stable training convergence, reaching a stable cooperative reward of 76.246 within approximately 371 training epochs. The framework further maintains near-linear computational scaling with the number of UAV agents, confirming its suitability for real-time deployment in dense, AI-native multi-UAV 6G wireless communication systems.

Read PDF

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.