Skip to content

Deep Reinforcement Learning for Autonomous Communication Networks: Resource Allocation, Spectrum Management, and Control

Jul 2026 · International Scientific Journal of Engineering and Management · Vol 05, pp. 1-9 · 0 citations

TL;DR

This paper (or study) explores the application of reinforcement learning techniques to autonomous communication networks, including resource allocation, spectrum management, power control, routing, and congestion control.

Abstract

ABSTRACT Autonomous communication systems are evolving toward self-organizing, adaptive networks capable of optimizing performance under dynamic and uncertain environments. Traditional rule-based and model-driven optimization techniques struggle to cope with the complexity, scale, and non-stationarity of modern wireless and networked systems. Reinforcement learning (RL), a branch of machine learning where agents learn optimal policies through interaction with the environment, has emerged as a powerful paradigm for enabling autonomy in communication systems. This paper (or study) explores the application of reinforcement learning techniques to autonomous communication networks, including resource allocation, spectrum management, power control, routing, and congestion control. By formulating communication tasks as Markov Decision Processes (MDPs), RL agents can learn to maximize long-term performance metrics such as throughput, latency, energy efficiency, and quality of service without requiring explicit mathematical models of the environment. Deep reinforcement learning (DRL), which integrates deep neural networks with RL, further enhances scalability by handling high-dimensional state and action spaces typical in modern networks such as 5G, 6G, and Internet of Things (IoT) systems. Multi-agent reinforcement learning (MARL) is also increasingly relevant, enabling distributed decision-making among multiple network nodes with partial observability and limited coordination. Despite its promise, RL-based communication systems face challenges including sample inefficiency, convergence stability, safety constraints, and real-time deployment limitations. Ongoing research focuses on improving training efficiency, incorporating domain knowledge, ensuring reliability, and developing hybrid models that combine RL with optimization and control theory. Overall, reinforcement learning provides a foundational framework for next-generation autonomous communication systems, enabling adaptive, intelligent, and self-optimizing networks. Keywords: Reinforcement Learning, Autonomous Communication Systems, Deep Reinforcement Learning, Multi-Agent Systems, Wireless Networks, Resource Allocation, Spectrum Management, Markov Decision Process, 5G/6G Networks, Internet of Things (IoT), Network Optimization, Self-Organizing Networks, Policy Learning, Dynamic Systems Optimization

View source

Similar papers

Book Open access Jul 2026

Deep Reinforcement Learning Driven Strategy Design for Distributed Consensus Black-Box Evolutionary Optimization

With the development of communication networks and smart terminals, distributed systems are crucial in critical domains like intelligent transportation and industrial internet. However, traditional distributed evolutionary algorithms struggle to address optimization demands in complex dynamic environments, as they are constrained by static parameter configurations and rigid communication topologies. To solve this, this paper introduces a novel mechanism where each agent autonomously learns its own evolutionary strategy via reinforcement learning (RL), enabling the automatic selection of learning targets and the adaptive adjustment of their weights. This mechanism is seamlessly integrated into the Multi-Agent Swarm Optimization with Internal and External Learning (MASOIE) framework, giving rise to RL-MASOIE. Experiments on heterogeneous benchmarks show RL-MASOIE outperforms the original MASOIE and existing black-box distributed algorithms in average fitness, with excellent robustness under diverse network topologies.

Duomao Zhuang, Wei-neng Chen, Tai-You Chen · 0 citations
Conference Jul 2026

Double Deep Reinforcement Learning–Based UAV Positioning for Throughput Optimization in Wireless Networks

This work investigates a reinforcement learning-based control framework for the autonomous movement and coordination of multiple Unmanned Aerial Vehicles (UAVs) in a wireless communication environment. The considered system includes UAVs performing sensing and relaying tasks, where mobility decisions directly affect the overall network performance. The main objective is to improve the communication quality of ground users by maximizing aggregate network throughput. To achieve this objective, a Double Deep Q-Network (DDQN) architecture is employed, where each UAV is assigned an individual learning agent. The agents learn role-specific movement policies while coordinating through interactions with the shared environment. Learning performance is further improved by using adaptive scaling and a custom reward function designed to capture variations in network utility. Simulation results show that the proposed approach outperforms baseline movement strategies in terms of utility. In addition, different task configurations, agent behaviors, and hyperparameter selections are examined to improve convergence speed and training stability. Overall, the results indicate that reinforcement learning is a promising method for cooperative UAV positioning in dynamic and interference-sensitive wireless communication scenarios.

Berke Kilinç, M. Ö. Efe · 0 citations
2026

Scalable Traffic Allocation in Dynamic Networks via End-to-End Imitation Learning

Networks with highly dynamic data transmission demands and network topologies are common in real world. A fundamental problem in such networks is achieving scalable traffic allocation to maximize long-term total throughput under link capacity constraints. However, state-of-the-art (SOTA) works lack scalability. This is primarily due to two reasons in large-scale networks: first, they require solving constrained optimization problems online, which leads to high decision latency; second, they rely on reinforcement learning algorithms for policy optimization, which are inefficient in exploration and challenging to train effectively. To address these issues, we propose the Fast Networked Control (FNC) policy framework, which firstly utilizes parallelizable neural network modules to process the state and generate raw decisions, followed by basic operations such as normalizations and comparisons, which do not require iteration or optimization, to obtain decisions that satisfy the constraints. Hence, FNC policy avoids solving constrained optimization problems and supports parallel execution, significantly reducing decision latency. Furthermore, this policy preserves gradient flow and supports backpropagation, which enable us to design an imitation learning algorithm to efficiently train the policy in an end-to-end manner. Experiments in large-scale networks show that our FNC policy achieves an average 8% improvement in demands satisfaction and 10 times reduction in decision latency versus SOTA works.

Zhaoxing Yang, Guiyun Fan, An-Jie Cao et al. · 0 citations
Open access Jul 2026

A Deep Reinforcement Learning Approach to Dynamic Channel Selection

Channel selection is a dynamic and challenging issue in the present wireless network that is shared and suffers interference in a spectrum that is not fully controlled. The traditional heuristic and greedy methods depend on immediate channel measurements and they are not always adapted to non-stationary environments resulting in poor throughput, high packet error rates and unpredictable channel switching behaviour. As a remedy to these shortcomings, this paper will suggest a deep reinforcement learning-based system to perform dynamic channel selection, which provides the formulation of the problem as a Markov Decision Process and uses a Deep Q-Network to train the optimal long-term policies of spectrum access. The given strategy combines spectrum sensing, contextual state representation, and reward-based learning to make adaptive and stable choices to select the channel. The reward function explicitly includes a switching cost in order to strike a balance between the adaptability and the stability. Numerous simulation-based tests have shown that the given technique replicates the best results in terms of mean throughput, packet error rate, and frequency of channel switching, compared to the random, greedy and bandit-based baseline schemes. The findings validate the outcomes of deep reinforcement learning in learning temporal spectrum dynamics and maximizing the long-term communication performance in non-stationary wireless conditions.

D. Shende, Dr. Amruta Nagesh Chitari, Dr. Pooja Mishra et al. · 0 citations
Open access Jul 2026

MULTI-AGENT REINFORCEMENT LEARNING FOR TASK OFFLOADING AND RESOURCE ALLOCATION IN MEC SYSTEMS

This paper addresses the joint task offloading and resource allocation problem in multi-user MEC systems and proposes a decentralized control framework based on Multi-Agent Reinforcement Learning (MARL), which achieves lower total system cost and faster convergence than the full-local, full-offload, and heuristic baselines.

Youssef Oukissou, Mohamed Amine Meddaoui, Ayoub Belaidi et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.