Aug 2026· Global Journal of Engineering and Technology Advances· 0 citations
TL;DR
A Deep Reinforcement Learning (DRL)-based framework for dynamic spectrum access in 6G heterogeneous Cognitive Radio Networks (Het-CRNs), wherein secondary users learn optimal channel selection policies through direct interaction with the radio environment, without requiring explicit statistical channel models is proposed.
Abstract
The proliferation of heterogeneous radio access technologies in sixth-generation (6G) wireless networks demands a fundamental rethinking of spectrum management strategies. Traditional spectrum sensing approaches, designed for relatively static channel conditions, are inadequate for the dynamic, interference-rich environments that characterize 6G deployments spanning sub-6 GHz, millimetre-wave, and terahertz bands simultaneously. This paper proposes a Deep Reinforcement Learning (DRL)-based framework for dynamic spectrum access in 6G heterogeneous Cognitive Radio Networks (Het-CRNs), wherein secondary users (SUs) learn optimal channel selection policies through direct interaction with the radio environment, without requiring explicit statistical channel models. Specifically, a Double Deep Q-Network (DDQN) architecture is adopted, augmented with a prioritized experience replay mechanism that accelerates policy convergence under non-stationary channel conditions. The proposed agent observes a composite state space encoding instantaneous channel occupancy, signal-to-interference-plus-noise ratio (SINR), primary user (PU) activity patterns, and residual energy levels, and selects actions that jointly optimize spectrum utilization efficiency, interference avoidance, and energy consumption. Simulation experiments conducted over a heterogeneous network topology with four primary users and eight secondary users demonstrate that the proposed DDQN-based scheme achieves a throughput gain of approximately 34% over conventional energy detection-based sensing, reduces interference to primary users by 61%, and attains a detection probability of 0.94 at a false alarm rate of 0.05. These results confirm the practical viability of DRL as a spectrum management backbone for next-generation cognitive radio systems.
Simulation results position D3QN-PER as a strong candidate for deployment as a near-RT RIC xApp within the O-RAN architecture, advancing the vision of AI-native mobility management for 6G.
Kalpesh Popat, Divyakant T. Meva· Telecommunications Systems· 0 citations
The ability of various isolated devices to sense their surroundings can be improved by 5G millimetre wave (mmWave) communication technology. By jointly supporting data transmission and sensing tasks, the framework improves overall spectrum efficiency in wireless networks. Among them, the Integrated Sensing and Communication (ISAC) has become the standard in wireless communications. Specifically, mmWave technology is highly effective for bandwidth-intensive communication services and delivers improved spatial and temporal accuracy through its large spectrum availability and directional beamforming characteristics. To meet the requirements, a multi-agent-based deep learning technique is proposed for better development. Over this sensing network of 5G mmWave, the resource allocation process is handled by Multi-agent Deep Reinforcement Learning with Prioritized Experience Replay (MDRL-PER), whereas the system is provided based on allocated resource for better communication. Finally, the performance of the system is assessed through distinct evaluation metrics and compared with existing methodologies. Hence, the superior results are obtained to ensure the efficacy of the communication network.
Papisetty Sai Prasad, T. Kavitha· 2026 7th International Confe...· 0 citations
An adaptive multi-mode Deep Reinforcement Learning (DRL) framework for intelligent RIS-assisted anti-jamming communication in dynamic 6G wireless networks that maintains stable communication performance under strong jamming power, CSI uncertainty, and high-mobility scenarios is proposed.
Lê Hoàng Hiệp, Huu-Huy Ngo· Journal of Communications So...· 1 citation
Sixth-generation (6G) mobile communication poses unprecedented challenges for resource scheduling under personalized demands. Cell-free massive multiple-input multiple-output (CF-mMIMO), with its user-centric characteristics, has emerged as a key technology for satisfying personalized demands. However, faced with heterogeneous quality-of-service (QoS) requirements, existing reinforcement learning schemes are constrained by partial observability, making it difficult to balance overall system performance and personalized demands. Consequently, we propose a graph-embedded multi-agent deep deterministic policy gradient (G-MADDPG) scheme. Guided by personalized demands, proposed G-MADDPG formulates a maximization problem for system weighted sum spectral efficiency and introduces differentiated QoS penalties. In addition, graph neural networks (GNNs) are embedded into the policy learning and value estimation processes of reinforcement learning, endowing agents with enhanced structural reception and cooperative capabilities. Simulation results demonstrate that proposed G-MADDPG scheme outperforms existing benchmark schemes in both convergence speed and performance evaluation.
Yu-Heng An· 2026 8th International Confe...· 0 citations
This work demonstrates the viability of RL for distributed resource management and provides a reproducible simulation toolkit to support further research in AI-driven wireless communication systems.
Mugerwa Joseph, Ajaegbu Chigozirim· International Journal Of Eng...· 0 citations
With the advent of seventh-generation (7G) wireless systems, the spectrum environment is extremely dynamic and heterogeneous, and traditional methods of sensing do not offer reliable and efficient performance. This paper introduces a self-evolving spectrum sensing system to enable adaptive and intelligent spectrum access based on a deep reinforcement-based learning paradigm. The framework depicts the sensing process as a sequence decision problem where an autonomous agent is able to continually refine its policy as it engages with the environment. The multi-objective reward formulation is designed to maximize the combination of the detection accuracy, false alarms, energy usage, and using the spectrum. Moreover, an adaptive representation of state mechanism is also introduced to indicate the temporal changes and short-term changes in the spectrum occupancy. The self-evolution strategy proposed adjusts learning parameters and decision policies in a dynamic manner that gives a robust operating in a non-stationary environment. The overall analysis of the experiment shows that the structure achieves detection probability of 97.1, false alarm rate reduced to minimum of 3.8, spectral usage maximized to above 92 and convergence rate is quicker as compared to the existing techniques. These results confirm the appropriateness of the proposed method in overcoming the issues of the next-generation wireless systems.
A.L Sriram, H. N. Divya, S. Sabarinathan et al.· International Conference on...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.