Skip to content
Conference

Feedback-Driven Online Deep Reinforcement Learning for Hotspot Aware Adaptive Traffic Steering in Non-Stationary Networks

Jul 2026 · 2026 6th International Conference on Inventive Computation and Information Technologies (ICICIT) · pp. 1630-1637 · 0 citations · 21 references

Abstract

Modern communication networks increasingly operate under non-stationary traffic conditions, where busty traffic and flash crowds challenge traditional static rule-based network control mechanisms. Despite the fact that reinforcement learning has already been explored for network optimization, most existing methods rely on offline-trained policies that lack stable adaptation to traffic in the network that causes high dimensional state. This paper proposes a feedback-driven online deep reinforcement learning framework for versatile traffic steering in mesh networks. The traffic steering problem is considered as a closed loop evaluation process in which a Deep Q-Network (DQN) constantly updates its policies during runtime. To balance the performance in the network, the framework implements a hotspot-aware lightweight state representation, composed of queue length and link utilization for the top three most congested links alongside end-to-end delay, packet loss rate, and throughput. The proposed framework achieves lower delay and packet loss, faster adaptation, and stable throughput compared to the existing static routing and offline-trained RL policies, while maintaining low monitoring overhead.

View source

Similar papers

Open access Jul 2026

STQ-Scheduler: A Secure and Throughput-Aware Deep Reinforcement Learning Framework for QoE-Driven Resource Scheduling in Distributed Video Streaming Systems

STQ-Scheduler is proposed, a secure and throughput-aware deep reinforcement learning framework that integrates high-throughput data processing, Transformer-based QoE prediction, and Proximal Policy Optimization-based resource scheduling to ensure data quality and prevent data processing from becoming a bottleneck in di...

Yi-Chun Chang, Min-Wei Jiang · 0 citations
2026

Scalable Traffic Allocation in Dynamic Networks via End-to-End Imitation Learning

Networks with highly dynamic data transmission demands and network topologies are common in real world. A fundamental problem in such networks is achieving scalable traffic allocation to maximize long-term total throughput under link capacity constraints. However, state-of-the-art (SOTA) works lack scalability. This is...

Zhaoxing Yang, Guiyun Fan, An-Jie Cao et al. · 0 citations
Open access Aug 2026

Dynamic Path Selection in SDN Based on Reinforcement Learning and Link Utilization

A path selection model that combines bottleneck link usage and reinforcement learning that achieves superior state awareness and adaptive routing performance in multi-source heterogeneous networks and hence can be used effectively for intelligent routing in next-generation power communication networks.

Ying Zeng, Xing-Nan Li, Yubeng Bao et al. · 0 citations
#artificial intelligence Preprint Sep 2026

From Prior-Guided Heuristics to Deployable Agents: Accelerating Demonstration-Driven Reinforcement Learning for Deadline-Constrained Network Control

This paper introduces Effective Congestion (EC), a deadline-aware metric family that quantifies interface congestion by packet urgency and proactively filters non-viable traffic, coupled with a Uniform Path Grouping (UPG) distribution heuristic promoting robust load-balancing; the resulting policies are embedded into M...

Vincenzo Norman Vitale, Mohammad Solki, A. Tulino et al. · 0 citations
Open access Aug 2026

RL-SDNTE: Reinforcement Learning-Driven Traffic Engineering in SDN for QoE Optimization in Video Streaming

RL-SDNTE is presented, a Reinforcement Learning-based TE framework built directly into an SDN controller that targets end-user Quality of Experience (QoE) as its primary objective and scales to topologies beyond 100 nodes without exceeding operationally acceptable convergence times.

Saurabh Suman, Roopali Lolag, Sanjay Sange et al. · 0 citations
Open access Jul 2026

Robust Offline Multi-Agent Reinforcement Learning for Latency-Aware SDN Path Control in 6G-Oriented Network Softwarization

Future sixth-generation (6G)-oriented networks require programmable control that can adapt routing to latency and congestion without unsafe online exploration. This study evaluates offline multi-agent deep deterministic policy gradient (MADDPG) with behavior-adjusted training rewards for latency-aware path control in s...

A. Kyzyrkanov, Y. Nurakhov, Zhenis Otarbay et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.