Jul 2026· International Conference on Ubiquitous and Future Networks· pp. 695-700· 0 citations· 14 references
Abstract
Ultra-dense 5G networks require advanced traffic steering to maintain performance and balance load amid growing user and base station (gNB) densities. Traditional heuristics such as nearest-base-station and Signal-to-Interference-plus-Noise Ratio (SINR)-based selection provide simple solutions but struggle to adapt to dynamic user mobility, diverse traffic, and fluctuating radio conditions at the mobility-control level, often leading to inefficient handovers and degraded network quality. We propose a deep reinforcement learning (DRL) framework to dynamically tune a global handover hysteresis margin that governs handover triggering decisions, optimizing handover success, reducing failures, and enhancing throughput and fairness. Implemented in Python using Stable Baselines3 and NumPy, our custom simulation environment models key mobility-related 5G dynamics at a high level, including user mobility, pathloss-based signal degradation, and interference. We evaluate DRL agents-Deep Q-Network (DQN) and Proximal Policy Optimization (PPO)-against heuristic and hysteresis-based baselines. Results show that DRL-based hysteresis optimization provides strong and robust performance under the considered ultra-dense mobility conditions in handover success rate, average SINR, throughput, and fairness, with PPO demonstrating the most consistent behavior across configurations. This work offers a reproducible simulation framework for further research into adaptive mobility management.
Simulation results position D3QN-PER as a strong candidate for deployment as a near-RT RIC xApp within the O-RAN architecture, advancing the vision of AI-native mobility management for 6G.
Kalpesh Popat, Divyakant T. Meva· Telecommunications Systems· 0 citations
Recently, the development and deployment of intelligent controllers for radio access networks (RAN) has attracted significant attention from network operators and international telecommunications organizations, driven by rapid advances in artificial intelligence. Mobility management plays a fundamental role in ensuring seamless connectivity and service quality in 5G RAN. In fact, optimal control in 5G RAN is highly challenging due to its complex, dynamic, and distributed environment. Many approaches have been proposed to address this problem, particularly those based on deep reinforcement learning (DRL). However, contrary to the dense reward assumption in many DRL-based studies, mobility feedback in practical RAN environments is characteristically sparse and delayed. In this paper, we propose WHO (World Model for Handover Optimization), a novel method designed to bridge the gap between sparse feedback and efficient learning in 5G networks. WHO utilizes a world model to convert event-driven rewards into dense predictive signals, facilitating robust multi-agent optimization. Field experiments involving 13 base stations and 39 cells show that the proposed method significantly improves handover performance and network stability compared to conventional DRL approaches, achieving 19–40% higher prediction precision and up to 32% improvement in key performance indicators (KPIs).
Uyen Thi Thu Truong, Doan Van Nguyen, Do Ngoc Tuan et al.· International Conference on...· 0 citations
A comprehensive survey of AI-enabled mobility management strategies for 5G, Beyond 5G, and upcoming 6G networks, with particular attention to HO optimization and load balancing is presented.
H. Asif, Abdulraqeb Alhammadi, N. Tarhuni et al.· Future Internet· 0 citations
This research is among the first to employ MARL to this extent, and it offers an end-to-end solution that combines cellular, Wi-Fi, and device-to-device (D2D) communications and considers practical network environments like user mobility and channel conditions.
Nabeel Abdolrazagh Yaseen Alrashedi, Rasool Sadeghi, Wael Hussein Zayer Al-Lamy et al.· Journal of universal compute...· 0 citations
A Hybrid deep double-Q networks (DDQN)–bidirectional long short-term memory (Bi-LSTM) Framework that integrates bi-directional mobility prediction and DRL-based adaptive decision-making is introduced that highlights the effectiveness of integrating predictive intelligence with reinforcement learning for reliable mobility management in 5G-Advanced and emerging 6G networks.
Target Wake Time (TWT) in IEEE 802.11ax enables significant power savings by coordinating station sleep schedules, but optimal TWT parameter selection under heterogeneous traffic remains an open challenge. We present a deep reinforcement learning framework for dynamic TWT scheduling that is protocol-compliant by design. The agent’s observation is restricted to 802.11ax buffer status reports, 802.11k station statistics, and AP-derived TWT metrics, ensuring direct deployment without firmware modifications. The system employs Proximal Policy Optimization (PPO) over a factored MultiDiscrete action space of 380 schedule-assignment combinations, trained end-to-end within ns-3 via shared-memory with episode-level process isolation for stable, reproducible training. We evaluate two architectures, MLP-PPO and LSTM-PPO, against an analytical M/D/1 baseline across three reward presets (energy, throughput, queue) over 50 seed-disjoint episodes with 16 stations from four heterogeneous device classes. Both RL agents substantially outperform the analytical baseline on all presets. On the energy preset, MLP-PPO reduces aggregate energy by 44% while improving throughput by 18% and reducing drops by 73%. On the queue preset, LSTM-PPO achieves 15% higher reward with 77% fewer drops by leveraging temporal correlations across beacon intervals. The two architectures exhibit complementary strengths. MLP-PPO excels on the energy-dominated preset where a stationary policy suffices while LSTM-PPO’s recurrent state captures multi-step queue-drain dynamics that transfer across network configurations, motivating future ensemble approaches.
A. Maksud, M. Carvalho· International Conference on...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.