Skip to content
Open access

RL-SDNTE: Reinforcement Learning-Driven Traffic Engineering in SDN for QoE Optimization in Video Streaming

Saurabh Suman Roopali Lolag Sanjay Sange Sonali Padalkar
Aug 2026 · International journal of computer information systems and industrial management applications · 0 citations

TL;DR

RL-SDNTE is presented, a Reinforcement Learning-based TE framework built directly into an SDN controller that targets end-user Quality of Experience (QoE) as its primary objective and scales to topologies beyond 100 nodes without exceeding operationally acceptable convergence times.

Abstract

Video streaming now accounts for over 80% of global Internet bandwidth, yet most SDN traffic engineering (TE) solutions still optimize for throughput and link utilization rather than what users actually experience. Poor startup times, frequent re-buffering, and unstable bit-rate remain common even on well-managed networks -- largely because the control plane has no visibility into application-layer quality. We present RL-SDNTE, a Reinforcement Learning-based TE framework built directly into an SDN controller that targets end-user Quality of Experience (QoE) as its primary objective. Rather than relying on a single proxy metric, RL-SDNTE feeds four perceptual indicators -- startup latency, re-buffering ratio, mean video quality, and bit-rate oscillation -- into a unified reward function that drives routing decisions. A Deep Q-Network (DQN) agent uses the controller’s global network view together with real-time client feedback to continuously adjust path selection. Testing on a Mini-net emulation platform and a physical 12-node SDN test-bed showed gains of up to 34% in composite QoE, 28% fewer re-buffering events, 22% lower startup latency, and 17% less quality oscillation compared to ECMP, OSPF, DEFO, and heuristic QoE- aware baselines [5]-[7],[15]. The system also scales to topologies beyond 100 nodes without exceeding operationally acceptable convergence times, making it viable for real-world SDN deployments.

Read PDF

Similar papers

2026

PPO-MS: Confidence-Aware and Collaborative Traffic Management for Multimedia Streaming in SDN

The rapid growth of multimedia streaming poses critical challenges, including bursty traffic and congestion, leading to playback delays. The existing separate prediction and control mechanisms for multimedia traffic scheduling, which are based on software-defined networks (SDN), are unable to proactively manage bursty traffic under uncertain conditions. This limitation is particularly evident in SDN-enabled backbone and multimedia-aware access networks, which typically assume centralized control and stable topologies. They lack integration of traffic prediction, traffic shaping, and real-time perception scheduling through reinforcement learning, resulting in low efficiency when exploring multiple paths in dynamic networks. To address this challenge, we propose PPO-MS (Proximal Policy Optimization-based Multimedia Scheduler), an SDN-based multimedia traffic scheduling algorithm integrating three key innovations: 1) A novel LSTM+HTB synergy where LSTM’s confidence intervals dynamically adjust HTB (Hierarchical Token Bucket) shaping parameters, enabling adaptive rate control under prediction uncertainty and overcoming the limitations of static LSTM+HTB hybrids; 2) A Deep Reinforcement Learning (DRL)-optimized path pruning method that reduces state and action spaces by generating a constrained set of $k$ disjoint candidate paths via an improved redundant tree algorithm. Unlike traditional multi-path schemes, this method tightly couples path preselection with the RL decision loop for adaptive, context-aware routing; 3) Generalized Advantage Estimation (GAE)-accelerated PPO for stable convergence in dynamic environments. In contrast to prior works (e.g., LSTM+RL for QoE or standalone tree algorithms), PPO-MS uniquely unifies these modules through confidence-aware traffic shaping and hierarchical decision-making, validated via comparative experiments. Results demonstrate that PPO-MS, through the synergistic integration of confidence-aware traffic shaping and DRL-optimized path pruning, significantly outperforms decoupled baselines. In particular, via isolation studies against simpler alternatives (e.g., mean-prediction and fixed-margin shaping), the confidence-aware shaping mechanism is validated to be superior under bursty traffic conditions. Overall, PPO-MS reduces end-to-end latency by 17.3% and packet loss by 32.4% while achieving 24.4% better load balancing during traffic bursts.

Jiawei Wu, Yibo Wang, Zelin Zhu · 0 citations
Open access Jul 2026

STQ-Scheduler: A Secure and Throughput-Aware Deep Reinforcement Learning Framework for QoE-Driven Resource Scheduling in Distributed Video Streaming Systems

STQ-Scheduler is proposed, a secure and throughput-aware deep reinforcement learning framework that integrates high-throughput data processing, Transformer-based QoE prediction, and Proximal Policy Optimization-based resource scheduling to ensure data quality and prevent data processing from becoming a bottleneck in distributed training.

Yi-Chun Chang, Min-Wei Jiang · 0 citations
Open access Jul 2026

Robust Offline Multi-Agent Reinforcement Learning for Latency-Aware SDN Path Control in 6G-Oriented Network Softwarization

Future sixth-generation (6G)-oriented networks require programmable control that can adapt routing to latency and congestion without unsafe online exploration. This study evaluates offline multi-agent deep deterministic policy gradient (MADDPG) with behavior-adjusted training rewards for latency-aware path control in software-defined networking (SDN). Each traffic pair is modeled as an agent selecting one of three retained candidate paths, while centralized critics learn coordinated decisions from topology-specific Ryu–Mininet transition datasets. Nine policies are compared using ten paired seeds on fat-tree, mesh-grid, and WAN-corridors topologies under a deployed utilization–latency weighting of 0.60/0.40, together with flow-completion, latency, congestion, architectural-comparison, sensitivity, robustness, statistical, and controller-overhead analyses. The utilization-aware path heuristic achieves the strongest overall reward ranking. MADDPG is the strongest learned policy on fat-tree, is not significantly outperformed by any evaluated policy on mesh-grid, and remains statistically tied with completion-matched policies on WAN-corridors. Behavior adjustment is topology-dependent rather than uniformly beneficial. The exported policy requires approximately 52μs per joint decision, whereas complete control-loop timing is dominated by network-statistics polling. These results support offline multi-agent SDN control as a competitive, low-overhead option when interpreted jointly with topology structure, flow completion, and strong heuristic baselines.

A. Kyzyrkanov, Y. Nurakhov, Zhenis Otarbay et al. · 0 citations
Open access Aug 2026

Prioritized Experience Replay-Based Deep Deterministic Policy Gradient for Reliable Path Selection in SDN-IoT Networks

: Routing optimization is becoming prominent in Software-Defined Networks (SDN) due to the exponential growth of network traffic demands and the requirement for Quality of Service (QoS). However, reliable routing that satisfies the QoS requirements, such as end-to-end delay, packet loss, and bandwidth, remains a difficult task in SDN. To overcome this limitation, a Deep Reinforcement Learning (DRL)-based Prioritized Experience Replay-based Deep Deterministic Policy Gradient (PER-DDPG) model is proposed to enhance the routing performance in SDN with Internet of Things (SDN-IoT) with QoS requirements. Initially, requests are received and prioritized using the postponement strategy technique in the SDN controller, and the weights of the links are evaluated using the DRL method. Then, a routing path is identified by the routing algorithm, and requests in the queue are released using the time-strategy technique. Hence, reliable routing in an SDN with QoS requirements is accomplished. The proposed routing model based on DRL is evaluated by utilizing the end-to-end latency, throughput, and packet loss.

Gaurav Kumar, G. Girisha, N. Shamanth · 0 citations
Review Open access 2026

Comprehensive Review of Optimization Techniques for User-Centric Distributed Network Slicing in 5G Networks

A QoE-aware framework for Multi-Access Edge Computing-enabled Open Radio Access Network (O-RAN) architectures, combining a graph attention network (GAT) encoder, distributed multi-agent DRL, and privacy-preserving FL, while transitioning control from Quality of Service (QoS) to QoE metrics is proposed.

Manoj Prasad Kunasegran, Wai Leong Pang, S. K. Phang · 0 citations
Conference Jul 2026

TS-D3QS: A Traffic-State-Aware Dueling Double-DQN Scheduler for Adaptive Network Queue Control

Adaptive queue management must balance throughput, delay, packet loss, and fairness under changing traffic and resource conditions. This paper proposes TS-D3QS, a traffic-state-aware queue scheduler that formulates multi-queue resource allocation as a discrete reinforcement-learning problem. The scheduler observes normalized queue and port features and selects one of 16 interpretable allocation profiles, each jointly specifying bandwidth shares, shared-buffer shares, and priority multipliers. A Dueling Double-DQN learner separates state value from action advantage and use a Double-DQN target to reduce value overestimation. Experiments in a reproducible four-queue simulator under bursty and non-stationary traffic show that TS-D3QS reduces latency by 7.54%, reduces aggregate loss by 3.21%, improves Jain fairness by 3.30%, and improves reward by 6.01% over vanilla DQN, while classical CoDel-like and WFQ-like rules remain competitive on selected objectives.

Hao-Yan Wang, Q. Guan, Dapeng Yan et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.