2026· IEEE Transactions on Network and Service Management· Vol 23, pp. 6836-6848· 0 citations· 41 references
Abstract
The rapid growth of multimedia streaming poses critical challenges, including bursty traffic and congestion, leading to playback delays. The existing separate prediction and control mechanisms for multimedia traffic scheduling, which are based on software-defined networks (SDN), are unable to proactively manage bursty traffic under uncertain conditions. This limitation is particularly evident in SDN-enabled backbone and multimedia-aware access networks, which typically assume centralized control and stable topologies. They lack integration of traffic prediction, traffic shaping, and real-time perception scheduling through reinforcement learning, resulting in low efficiency when exploring multiple paths in dynamic networks. To address this challenge, we propose PPO-MS (Proximal Policy Optimization-based Multimedia Scheduler), an SDN-based multimedia traffic scheduling algorithm integrating three key innovations: 1) A novel LSTM+HTB synergy where LSTM’s confidence intervals dynamically adjust HTB (Hierarchical Token Bucket) shaping parameters, enabling adaptive rate control under prediction uncertainty and overcoming the limitations of static LSTM+HTB hybrids; 2) A Deep Reinforcement Learning (DRL)-optimized path pruning method that reduces state and action spaces by generating a constrained set of $k$ disjoint candidate paths via an improved redundant tree algorithm. Unlike traditional multi-path schemes, this method tightly couples path preselection with the RL decision loop for adaptive, context-aware routing; 3) Generalized Advantage Estimation (GAE)-accelerated PPO for stable convergence in dynamic environments. In contrast to prior works (e.g., LSTM+RL for QoE or standalone tree algorithms), PPO-MS uniquely unifies these modules through confidence-aware traffic shaping and hierarchical decision-making, validated via comparative experiments. Results demonstrate that PPO-MS, through the synergistic integration of confidence-aware traffic shaping and DRL-optimized path pruning, significantly outperforms decoupled baselines. In particular, via isolation studies against simpler alternatives (e.g., mean-prediction and fixed-margin shaping), the confidence-aware shaping mechanism is validated to be superior under bursty traffic conditions. Overall, PPO-MS reduces end-to-end latency by 17.3% and packet loss by 32.4% while achieving 24.4% better load balancing during traffic bursts.
The results indicate that iScavenger provides configurable operating points in the latency–utilization trade-off, limiting additional Sticky-flow RTT while achieving higher background throughput than conservative baseline policies, and highlight the potential of short-term traffic-demand prediction for proactive contention management in ATSSS-enabled multi-access networks.
Shah M. Emad Uddin, Karl-Johan Grinnemo, Arunselvan Ramaswamy et al.· IEEE Open Journal of the Com...· 0 citations
RL-SDNTE is presented, a Reinforcement Learning-based TE framework built directly into an SDN controller that targets end-user Quality of Experience (QoE) as its primary objective and scales to topologies beyond 100 nodes without exceeding operationally acceptable convergence times.
Saurabh Suman, Roopali Lolag, Sanjay Sange et al.· International journal of com...· 0 citations
A hybrid reinforcement learning (RL) framework that jointly controls queue management and bandwidth allocation in bursty multi-service networks and demonstrates the effectiveness of coordinated learning-based control for stable and QoS-aware operation in bursty networked systems.
T. Khan, Babar Shah, Taimur Karamat et al.· Computing· 0 citations
A path selection model that combines bottleneck link usage and reinforcement learning that achieves superior state awareness and adaptive routing performance in multi-source heterogeneous networks and hence can be used effectively for intelligent routing in next-generation power communication networks.
Ying Zeng, Xingnan Li, Yubeng Bao et al.· EAI Endorsed Transactions on...· 0 citations
The rapid evolution of 5G and emerging 6G networks requires optical access systems to support immersive extended reality (XR) services with stringent quality-of-service (QoS) requirements, like ultra-low latency and high bandwidth. However, conventional dynamic bandwidth allocation (DBA) schemes in passive optical networks (PONs) allocate upstream bandwidth solely based on reported queue occupancy, without considering the unique characteristics of XR traffic. To address these limitations, we propose an XR-aware Predictive (XP)-DBA scheme that integrates XR traffic prediction, deadline-aware scheduling, adaptive grant control, and a cycle-controller to proactively allocate bandwidth, prioritize latency-critical packets, and limit polling-cycle growth. We also derive closed-form analytical expressions to characterize XR-specific stability and delay feasibility in PON systems. We evaluate XP-DBA under standardized and burst-enhanced XR traffic models across varying XR user densities and transmission distances of up to 100 km. The results show that XP-DBA will reduce latency, jitter, and polling-cycle time while increasing throughput and supporting higher XR user densities under heavy network loads without violating XR delay bounds. These findings establish XP-DBA as an efficient and scalable scheduling solution for next-generation immersive XR services over long-reach optical access networks.
Akhilesh Patel, Y. Singh· IEEE Transactions on Network...· 0 citations