Skip to content
Preprint

Deep Reinforcement Learning Orchestration of Game-Theoretic User Association and Resource Allocation in HetNets

Aug 2026 · 0 citations · 32 references
Engineering

TL;DR

A novel orchestration scheme for game-theoretic UARA in HetNets that closely approximates the optimal policy for the considered operational objectives, while delivering higher network throughput than conventional association methods.

Abstract

Managing dynamic User Association and Resource Allocation (UARA) in modern Heterogeneous Cellular Networks (HetNets) remains a critical open challenge. Existing mathematical optimization and Reinforcement Learning approaches face limitations in handling low-latency decision-making under dynamic traffic conditions. This paper introduces a novel orchestration scheme for game-theoretic UARA in HetNets. The proposed bilevel framework distributes UARA decisions to User Equipment through a multi-objective non-cooperative game. Overlaying the distributed game, a centralized Deep Reinforcement Learning controller orchestrates network performance by dynamically configuring the game's utility parameters, enabling transitions between power awareness, coverage enhancement, and balanced operation. Evaluated on urban HetNet topologies with 3GPP TR 38.901-compliant channel modeling, the proposed framework closely approximates the optimal policy for the considered operational objectives, while delivering higher network throughput than conventional association methods. Furthermore, it incurs low computational overhead and maintains stable performance across the evaluated traffic densities without retraining.

View source

Similar papers

Open access Aug 2026

JATO: Deep Reinforcement Learning-based Joint Optimization for Task Offloading and Adaptive Transmission in Multimedia IoT Systems

As the Multimedia Internet of Things (M-IoT) evolves, the orchestration of numerous resources that offer support for high-bandwidth, low-latency applications arises as a key challenge. Architecturally, the edge-cloud framework alleviates structural concerns, but the linked nature of compute and data transfer poses problems of resource management. Approaches that tackle task offloading and adaptive transmission that think independently of each other tend to have problems such as user-server cross-region overloads or network congestion. This paper presents JATO, a framework to jointly tackle the problems of adaptive task offloading and transmission optimization using Deep Reinforcement Learning. JATO offers a mono-faceted solution, learning a policy to simultaneously determine the best offloading target and the transmission quality. The framework was implemented for evaluation with a combination of different edge devices in a testbed alongside a simulation environment. JATO recorded a result of 0.9321 as the holistic score of the overall framework endpoint, a score significantly better than that of all the other frameworks that were used as functional baselines. JATO was able to resource optimally with a network lag of 131.65 milliseconds and a network freeze of 0.09% with the resources utilized. This is evidence that offloading and rate control in combination provides better resource elasticity for M-IoT systems.

G. Purnama, Irma Amelia Dewi, A. Langi et al. · 0 citations
Open access Jul 2026

Robust Offline Multi-Agent Reinforcement Learning for Latency-Aware SDN Path Control in 6G-Oriented Network Softwarization

Future sixth-generation (6G)-oriented networks require programmable control that can adapt routing to latency and congestion without unsafe online exploration. This study evaluates offline multi-agent deep deterministic policy gradient (MADDPG) with behavior-adjusted training rewards for latency-aware path control in software-defined networking (SDN). Each traffic pair is modeled as an agent selecting one of three retained candidate paths, while centralized critics learn coordinated decisions from topology-specific Ryu–Mininet transition datasets. Nine policies are compared using ten paired seeds on fat-tree, mesh-grid, and WAN-corridors topologies under a deployed utilization–latency weighting of 0.60/0.40, together with flow-completion, latency, congestion, architectural-comparison, sensitivity, robustness, statistical, and controller-overhead analyses. The utilization-aware path heuristic achieves the strongest overall reward ranking. MADDPG is the strongest learned policy on fat-tree, is not significantly outperformed by any evaluated policy on mesh-grid, and remains statistically tied with completion-matched policies on WAN-corridors. Behavior adjustment is topology-dependent rather than uniformly beneficial. The exported policy requires approximately 52μs per joint decision, whereas complete control-loop timing is dominated by network-statistics polling. These results support offline multi-agent SDN control as a competitive, low-overhead option when interpreted jointly with topology structure, flow completion, and strong heuristic baselines.

A. Kyzyrkanov, Y. Nurakhov, Zhenis Otarbay et al. · 0 citations
Open access Jul 2026

MULTI-AGENT REINFORCEMENT LEARNING FOR TASK OFFLOADING AND RESOURCE ALLOCATION IN MEC SYSTEMS

This paper addresses the joint task offloading and resource allocation problem in multi-user MEC systems and proposes a decentralized control framework based on Multi-Agent Reinforcement Learning (MARL), which achieves lower total system cost and faster convergence than the full-local, full-offload, and heuristic baselines.

Youssef Oukissou, Mohamed Amine Meddaoui, Ayoub Belaidi et al. · 0 citations
Open access 2026

Scalable Reinforcement Learning-Based Environment-Aware Scheduler for Enhancing Reliability and Availability in 6G Networks

A scalable deep reinforcement learning (DRL) framework that exploits environmental-aware knowledge to optimize multi-user scheduling under limited resources, in order to enhance reliability and availability while maintaining fairness, outperforming Round Robin and Proportional Fair schedulers.

Roya Khanzadeh, Fjolla Ademaj-Berisha, Bernhard Etzlinger et al. · 0 citations
Conference Jul 2026

AP Selection and Power Control for Personalized Cell-Free Massive MIMO: Graph-Embedded Reinforcement Learning Approach

Sixth-generation (6G) mobile communication poses unprecedented challenges for resource scheduling under personalized demands. Cell-free massive multiple-input multiple-output (CF-mMIMO), with its user-centric characteristics, has emerged as a key technology for satisfying personalized demands. However, faced with heterogeneous quality-of-service (QoS) requirements, existing reinforcement learning schemes are constrained by partial observability, making it difficult to balance overall system performance and personalized demands. Consequently, we propose a graph-embedded multi-agent deep deterministic policy gradient (G-MADDPG) scheme. Guided by personalized demands, proposed G-MADDPG formulates a maximization problem for system weighted sum spectral efficiency and introduces differentiated QoS penalties. In addition, graph neural networks (GNNs) are embedded into the policy learning and value estimation processes of reinforcement learning, endowing agents with enhanced structural reception and cooperative capabilities. Simulation results demonstrate that proposed G-MADDPG scheme outperforms existing benchmark schemes in both convergence speed and performance evaluation.

Yu-Heng An · 0 citations
Open access Jul 2026

Generalised Potential Game-Based Resource Allocation in SDN-Enabled O-RAN Systems

This formulation provides a rigorous and tractable framework for distributed spectrum sharing in 6G O-RAN systems, with the potential to support intelligent and adaptive control in future wireless networks.

E. Spyrou, Chrysostomos D. Stylios, V. Kappatos et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.