Skip to content
Open access

Hybrid Deep Reinforcement Learning with Kookaburra Optimization for QOS-Aware Bandwidth Allocation in GMPLS Optical Networks

Aug 2026 · Indian Journal of Science and Technology · Vol 19, pp. 2140-2151 · 0 citations

TL;DR

The simulation results show that the suggested Hyb-DRL-KkOA algorithm performs better than the conventional bandwidth allocation algorithms, and reduces the blocking probability, makespan, cost, and energy utilization while improving the throughput; therefore, providing enhanced quality of service (QoS).

Abstract

Objectives: To address dynamic bandwidth allocation with strict Quality of Service (QoS) requirements in Generalized Multi-Protocol Label Switching (GMPLS) optical networks under strain from internet services, real-time multimedia, and cloud infrastructure. Method: A Hybrid Deep Reinforcement Learning (Hyb-DRL) framework combined with the Kookaburra Optimization Algorithm (KkOA) for adaptive weight adjustment is proposed. Dynamically generated input data, including user request rates, queue lengths, server availability, and link stability metrics, were used to simulate real-world traffic. The Hyb-DRL agent learned optimal routing and bandwidth provisioning policies while KkOA optimized model weights for faster convergence and stability. Findings: The simulation results show that the suggested Hyb-DRL-KkOA algorithm performs better than the conventional bandwidth allocation algorithms. In contrast to conventional algorithms, it reduces the blocking probability, makespan, cost, and energy utilization while improving the throughput; therefore, providing enhanced quality of service (QoS). The proposed framework achieves a lower blocking probability by 78%, makespan by 64%, energy consumption by 51%, and operational cost by 47%. In addition to this, it provides better throughput performance by 69% than other conventional techniques. Moreover, it provided a delay of 0.0189 s, minimal energy consumption of 33 mJ, and maximal throughput of 950 Mbps. Novelty: A combination of reinforcement learning and meta-heuristic optimization leads to adaptive decision-making regarding routing and bandwidth allocation in the face of different traffic demands. The performance gain in terms of QoS is due to optimal utilization of network resources with low blocking probability, energy and operational cost. It provides a scalable and adaptive solution for high-speed, reliable data transmission in modern communication networks. Keywords: GMPLS Optical Networks, Kookaburra Optimization Algorithm, Bandwidth Allocation, Quality of Service (QoS), Blocking Probability

Read PDF

Similar papers

Open access Jul 2026

Deep Reinforcement Learning-Based QoS-Aware Routing Protocol for Space–Air–Ground Integrated Networks

A deep reinforcement learning (DRL)-based adaptive routing scheme for maximizing throughput and minimizing end-to-end delay jointly in SAGIN and indicates that adaptive policy learning enables better congestion avoidance and more efficient resource utilization.

Nilu Mishra, Sanakat Bhanjan Prusty, Sachin Sharma · 0 citations
Open access Aug 2026

Prioritized Experience Replay-Based Deep Deterministic Policy Gradient for Reliable Path Selection in SDN-IoT Networks

: Routing optimization is becoming prominent in Software-Defined Networks (SDN) due to the exponential growth of network traffic demands and the requirement for Quality of Service (QoS). However, reliable routing that satisfies the QoS requirements, such as end-to-end delay, packet loss, and bandwidth, remains a difficult task in SDN. To overcome this limitation, a Deep Reinforcement Learning (DRL)-based Prioritized Experience Replay-based Deep Deterministic Policy Gradient (PER-DDPG) model is proposed to enhance the routing performance in SDN with Internet of Things (SDN-IoT) with QoS requirements. Initially, requests are received and prioritized using the postponement strategy technique in the SDN controller, and the weights of the links are evaluated using the DRL method. Then, a routing path is identified by the routing algorithm, and requests in the queue are released using the time-strategy technique. Hence, reliable routing in an SDN with QoS requirements is accomplished. The proposed routing model based on DRL is evaluated by utilizing the end-to-end latency, throughput, and packet loss.

Gaurav Kumar, G. Girisha, N. Shamanth · 0 citations
Open access Aug 2026

JATO: Deep Reinforcement Learning-based Joint Optimization for Task Offloading and Adaptive Transmission in Multimedia IoT Systems

As the Multimedia Internet of Things (M-IoT) evolves, the orchestration of numerous resources that offer support for high-bandwidth, low-latency applications arises as a key challenge. Architecturally, the edge-cloud framework alleviates structural concerns, but the linked nature of compute and data transfer poses problems of resource management. Approaches that tackle task offloading and adaptive transmission that think independently of each other tend to have problems such as user-server cross-region overloads or network congestion. This paper presents JATO, a framework to jointly tackle the problems of adaptive task offloading and transmission optimization using Deep Reinforcement Learning. JATO offers a mono-faceted solution, learning a policy to simultaneously determine the best offloading target and the transmission quality. The framework was implemented for evaluation with a combination of different edge devices in a testbed alongside a simulation environment. JATO recorded a result of 0.9321 as the holistic score of the overall framework endpoint, a score significantly better than that of all the other frameworks that were used as functional baselines. JATO was able to resource optimally with a network lag of 131.65 milliseconds and a network freeze of 0.09% with the resources utilized. This is evidence that offloading and rate control in combination provides better resource elasticity for M-IoT systems.

G. Purnama, Irma Amelia Dewi, A. Langi et al. · 0 citations
Open access Aug 2026

Dynamic Path Selection in SDN Based on Reinforcement Learning and Link Utilization

A path selection model that combines bottleneck link usage and reinforcement learning that achieves superior state awareness and adaptive routing performance in multi-source heterogeneous networks and hence can be used effectively for intelligent routing in next-generation power communication networks.

Ying Zeng, Xingnan Li, Yubeng Bao et al. · 0 citations
Jul 2026

Deep reinforcement learning-based multi-objective routing and resource allocation strategy in quantum key distribution optical networks

Quantum key distribution (QKD) optical networks, relying on the information-theoretic security guaranteed by the fundamental laws of quantum physics, have become a core solution for ensuring data transmission security in critical fields such as finance and energy. However, QKD optical networks have inherent characteristics like severe resource constraints, dynamic changes in network environments. The collaborative optimization of routing, wavelength, and time-slot assignment (RWTA) has thus become a key bottleneck restricting their practical deployment. Traditional algorithms suffer from high complexity and poor adaptability in large-scale dynamic networks, while existing deep reinforcement learning (DRL) schemes insufficiently consider network security states, failing multi-objective co-optimization. Therefore, this paper proposes a multi-objective secure proximal policy optimization algorithm (MOS-PPO) based on DRL to solve the routing, wavelength, and time slots assignment problem in QKD optical networks. The MOS-PPO algorithm expands the state space by introducing a network security state vector that includes quantum bit error rate, risk coefficient, node capacity, and processing capability. It designs a multi-objective reward function, and establishes a security constraint mechanism to achieve the collaborative optimization of blocking probability (BP), resource utilization (RU), and security performance. Simulations on NSFNET and UBN24 topologies show MOS-PPO outperforms DRL-PPO algorithm, Security-aware Load balancing, first fit algorithm, and random fit algorithm. On NSFNET, average BP drops by 20.4%, 22.1%,28.4%, 36.7% and RU rises by 4.0%,3.9%, 22.2%, 31.0%, on UBN24, BP decreases by 9.3%, 11.4%, 18.8%, and 27.4% and RU increases by 3.8%, 21.2%, 10.3%, and 29.5%, with higher stable security margins. This study verifies the high adaptability and multi-objective optimization ability of the MOS-PPO algorithm in dynamic network environments, providing an efficient RWTA solution for the practical deployment of QKD optical networks.

Li Liu, Zhenjiang Feng, Wenbin Ke et al. · 0 citations
Conference Jul 2026

AP Selection and Power Control for Personalized Cell-Free Massive MIMO: Graph-Embedded Reinforcement Learning Approach

Sixth-generation (6G) mobile communication poses unprecedented challenges for resource scheduling under personalized demands. Cell-free massive multiple-input multiple-output (CF-mMIMO), with its user-centric characteristics, has emerged as a key technology for satisfying personalized demands. However, faced with heterogeneous quality-of-service (QoS) requirements, existing reinforcement learning schemes are constrained by partial observability, making it difficult to balance overall system performance and personalized demands. Consequently, we propose a graph-embedded multi-agent deep deterministic policy gradient (G-MADDPG) scheme. Guided by personalized demands, proposed G-MADDPG formulates a maximization problem for system weighted sum spectral efficiency and introduces differentiated QoS penalties. In addition, graph neural networks (GNNs) are embedded into the policy learning and value estimation processes of reinforcement learning, endowing agents with enhanced structural reception and cooperative capabilities. Simulation results demonstrate that proposed G-MADDPG scheme outperforms existing benchmark schemes in both convergence speed and performance evaluation.

Yu-Heng An · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.