Skip to content
Conference

Reinforcement Learning for Efficient Link Scheduling in Multi-Link Networks

Jul 2026 · International Mediterranean Conference on Communications and Networking · pp. 1-6 · 0 citations · 11 references
Computer Science

Abstract

The increasing demand for higher throughput and lower latency in modern applications has driven the evolution of IEEE 802.11 with the introduction of Wi-Fi 7. The new Extremely High Throughput (EHT) amendment enhances performance with Multi-Link Operation (MLO), allowing concurrent transmissions over multiple frequency links. While MLO improves channel access, reduces latency, and boosts reliability, uneven traffic loads may still cause link congestion, starvation, or collisions, which can severely impact TCP flows. This work investigates the impact of MLO on TCP best effort and video traffic flows, and proposes a Reinforcement Learning (RL) Transmission Opportunity (TXOP) Random Discard strategy for adaptive load balancing. Simulation results demonstrate that the proposed AI-driven approach enhances TCP throughput and latency in Wi-Fi 7 multi-link networks.

View source

Similar papers

Open access Aug 2026

Dynamic Path Selection in SDN Based on Reinforcement Learning and Link Utilization

A path selection model that combines bottleneck link usage and reinforcement learning that achieves superior state awareness and adaptive routing performance in multi-source heterogeneous networks and hence can be used effectively for intelligent routing in next-generation power communication networks.

Ying Zeng, Xingnan Li, Yubeng Bao et al. · 0 citations
Conference Jul 2026

Adaptive Scheduling for Low-Latency Coordination in Wi-Fi 8 Multi-AP Networks

Reliable low-latency communication is a critical requirement in enterprise wireless networks such as hospitals, offices, and campuses. This paper proposes an earliest deadline first (EDF)-Lyapunov-Robbins-Monro (ELR), a stochastic scheduling algorithm for IEEE 802.11bn (Wi-Fi 8) Multi-Access Point Coordination Coordinated-Spatial Reuse (MAPC C-SR) networks that jointly accounts for queue stability and deadline-aware latency regulation under bursty traffic. A Lyapunov drift-based criterion for a group is adopted to ensure queues remain stable under varying traffic loads. Since the optimal balance between queue backlog and deadline urgency cannot be determined a priori under bursty traffic, EDF term is incorporated into the selection metric with a tunable balance parameter $\alpha$, governed by Robbins-Monro stochastic approximation scheme. The proposed algorithm addresses the inability of existing schedulers to track sudden congestion under bursty traffic, by dynamically adjusting $\alpha$ to suppress sharp delay spikes. Simulations over a four-access point (AP) enterprise deployment under bursty Markov-Modulated Poisson Process (MMPP) traffic demonstrate that ELR achieves 14.23%, 13.26%, and 7.97% reduction in 99th percentile delay over maximum number of packets (MNP), oldest packet (OP), and traffic alignment tracker (TAT) respectively under high load with 16 stations (STAs).

Hiya Shah · 0 citations
Open access Aug 2026

Hybrid DQN–PPO control for joint queue management and bandwidth allocation under bursty network traffic

A hybrid reinforcement learning (RL) framework that jointly controls queue management and bandwidth allocation in bursty multi-service networks and demonstrates the effectiveness of coordinated learning-based control for stable and QoS-aware operation in bursty networked systems.

T. Khan, Babar Shah, Taimur Karamat et al. · 0 citations
2026

PPO-MS: Confidence-Aware and Collaborative Traffic Management for Multimedia Streaming in SDN

The rapid growth of multimedia streaming poses critical challenges, including bursty traffic and congestion, leading to playback delays. The existing separate prediction and control mechanisms for multimedia traffic scheduling, which are based on software-defined networks (SDN), are unable to proactively manage bursty traffic under uncertain conditions. This limitation is particularly evident in SDN-enabled backbone and multimedia-aware access networks, which typically assume centralized control and stable topologies. They lack integration of traffic prediction, traffic shaping, and real-time perception scheduling through reinforcement learning, resulting in low efficiency when exploring multiple paths in dynamic networks. To address this challenge, we propose PPO-MS (Proximal Policy Optimization-based Multimedia Scheduler), an SDN-based multimedia traffic scheduling algorithm integrating three key innovations: 1) A novel LSTM+HTB synergy where LSTM’s confidence intervals dynamically adjust HTB (Hierarchical Token Bucket) shaping parameters, enabling adaptive rate control under prediction uncertainty and overcoming the limitations of static LSTM+HTB hybrids; 2) A Deep Reinforcement Learning (DRL)-optimized path pruning method that reduces state and action spaces by generating a constrained set of $k$ disjoint candidate paths via an improved redundant tree algorithm. Unlike traditional multi-path schemes, this method tightly couples path preselection with the RL decision loop for adaptive, context-aware routing; 3) Generalized Advantage Estimation (GAE)-accelerated PPO for stable convergence in dynamic environments. In contrast to prior works (e.g., LSTM+RL for QoE or standalone tree algorithms), PPO-MS uniquely unifies these modules through confidence-aware traffic shaping and hierarchical decision-making, validated via comparative experiments. Results demonstrate that PPO-MS, through the synergistic integration of confidence-aware traffic shaping and DRL-optimized path pruning, significantly outperforms decoupled baselines. In particular, via isolation studies against simpler alternatives (e.g., mean-prediction and fixed-margin shaping), the confidence-aware shaping mechanism is validated to be superior under bursty traffic conditions. Overall, PPO-MS reduces end-to-end latency by 17.3% and packet loss by 32.4% while achieving 24.4% better load balancing during traffic bursts.

Jiawei Wu, Yibo Wang, Zelin Zhu · 0 citations
Open access Aug 2026

Energy-Efficient Cooperative Data Offloading in Cellular Networks Using Reinforcement Learning

This research is among the first to employ MARL to this extent, and it offers an end-to-end solution that combines cellular, Wi-Fi, and device-to-device (D2D) communications and considers practical network environments like user mobility and channel conditions.

Nabeel Abdolrazagh Yaseen Alrashedi, Rasool Sadeghi, Wael Hussein Zayer Al-Lamy et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.