Aug 2026· Transactions on Emerging Telecommunications Technologies· 0 citations· 26 references
TL;DR
A Quantum Federated Reinforcement Learning (QFRL)‐based traffic offloading framework for RSMA‐enabled SAGINs is proposed, allowing distributed small cells to jointly optimize traffic offloading ratios, bandwidth allocation, RSMA power distribution, and UAV trajectory planning while satisfying stringent delay and reliability requirements.
Abstract
The evolution of sixth‐generation (6G) wireless networks demands ultra‐reliable low‐latency communication (URLLC), massive connectivity, and high‐capacity data transmission in highly dynamic environments. Space–Air–Ground Integrated Networks (SAGINs) have emerged as a promising architecture by seamlessly integrating satellites, unmanned aerial vehicles (UAVs), and terrestrial infrastructure to provide ubiquitous connectivity. However, stochastic traffic arrivals, UAV mobility, time‐varying wireless channels, and the coexistence of enhanced Mobile Broadband (eMBB) and URLLC services make traffic offloading and resource allocation highly challenging. These factors transform the optimization task into a stochastic mixed‐integer nonlinear programming (MINLP) problem. To address this challenge, this paper proposes a Quantum Federated Reinforcement Learning (QFRL)‐based traffic offloading framework for RSMA‐enabled SAGINs. The optimization problem is formulated as a constrained Markov decision process (CMDP), allowing distributed small cells to jointly optimize traffic offloading ratios, bandwidth allocation, RSMA power distribution, and UAV trajectory planning while satisfying stringent delay and reliability requirements. A variational quantum circuit (VQC)‐based actor‐critic architecture is developed to improve learning efficiency and policy representation in high‐dimensional continuous action spaces. In addition, a federated aggregation mechanism enables privacy‐preserving distributed learning and scalable coordination across the space, air, and ground segments. The proposed framework employs temporal‐difference learning and parameter‐shift gradient optimization to ensure stable convergence under stochastic network dynamics. Simulation results demonstrated that the proposed QFRL framework reduces traffic dropping probability by 28%–35%, decreases URLLC delay by 22%–30%, improves network availability by 18%–25%, and enhances traffic offloading efficiency by 20%–27% compared with Differentiated Federated Soft Actor‐Critic (DFSAC), Double Q‐Learning delay sensitive replay memory (DSRPM), and Nash Equilibrium Iteration Offloading (NEIO‐G) schemes, respectively.
Unmanned Aerial Vehicles (UAVs) have emerged as a flexible, cost-effective solution for connecting Internet of Things (IoT) devices where traditional infrastructure falls short. However, managing their limited energy alongside the diverse demands of densely deployed devices makes resource allocation a genuinely hard problem. This paper presents a Deep Reinforcement Learning (DRL) framework that jointly optimizes user scheduling, IoT device transmit power, bandwidth, and UAV movement in a 6G-enabled UAV-relay uplink network, using a deterministic large-scale air-to-ground path-loss channel model. The UAV acts as an aerial decode-and-forward relay between IoT devices and a Base Station (BS), with a Deep Q-Network (DQN) making decisions based on queue backlogs, channel conditions, UAV position, and remaining battery. The reward function balances Energy Efficiency (EE), queue stability, fairness, and battery longevity. We benchmark the DQN against six baselines; Round Robin (RR), Random Allocation (RA), the Single-to-Noise Ratio (Max-SNR), Proportional Fair (PF), a Lyapunov heuristic, and a GreedyEE scheme; across a range of device counts, traffic loads, battery budgets, and flight altitudes. Simulations consistently show that the DQN outperforms all baselines, including a RA baseline with equal access to UAV mobility; in EE, throughput, delay, and fairness, confirming that the gain stems from the learned joint control policy rather than from UAV mobility being available.
Alissa Nauman, Sung Won Kim· Italian National Conference...· 0 citations
: Space-Air-Ground Integrated Networks (SAGIN) provide a multi-layered, wide-coverage computing infrastructure for distributed urban sensing systems. However, their heterogeneity and dynamics pose unprecedented challenges for task offloading and resource allocation. Existing methods struggle to simultaneously address the complexity of cross-layer decision-making and reliability assurance under uncertain conditions. This paper proposes a novel framework, termed DRL-RA, which synergistically integrates Deep Reinforcement Learning (DRL) with reliability-aware optimization. The framework consists of two complementary components: (1) a Dueling Double Deep Q-Network (D3QN) module that learns adaptive policies to make offloading decisions among various options including local execution, terrestrial edge, UAVs, and satellites; (2) a Reliability-Aware Multi-Objective Optimization Framework (RA-MOOF) that introduces explicit reliability guarantees through cross-layer link reliability modeling, node availability estimation, and smooth reliability proxy functions. Addressing the heterogeneous communication characteristics of the SAGIN architecture, this paper establishes a complete cross-layer delay model and composite reliability metrics. The reliability formulation is defined under explicitly stated conditional-independence assumptions, and the proposed smooth constraint terms are treated as surrogate CMDP costs rather than exact hard chance-constraint guarantees. Extensive experiments in a SAGIN simulation environment demonstrate that the proposed method improves the task completion rate by 3.8%, reduces average latency by 11.1%, and increases system reliability by 3.9% compared to state-of-the-art benchmarks. The optimization-only RA-Opt baseline is used as a non-real-time optimization reference for assessing reliability-aware offloading decision quality, while deployment-time decision-latency comparisons are interpreted primarily among learned inference policies. Comprehensive ablation studies and statistical validation across multiple random seeds confirm the contributions of each component, while cross-layer offloading decision analysis verifies the effectiveness of the method across different network layer selections.
Fei-Yan Bu, Zheng Wang, Yong Pan et al.· Computers, Materials & C...· 0 citations
AF-EdgeRL is proposed, a novel Byzantine-resilient Asynchronous Federated Reinforcement Learning framework tailored for distributed resource allocation and dynamic task offloading and establishes theoretical convergence guarantees under non-convex reinforcement learning objectives.
Daniel Merrow, Tember L. Nair, Lucas Farnandez· International Journal of App...· 0 citations
Results show that topology‐aware cooperative learning can provide a scalable and practical solution for reliable UAV communication in future intelligent aerial and 6G‐enabled networks, particularly in dense deployments.
V. Nam, A. Chehri, Weiwei Jiang et al.· Expert Syst. J. Knowl. Eng.· 0 citations
Low Earth Orbit (LEO) constellations are required to process increasing volumes of heterogeneous tasks from ground networks. Intermittent inter-satellite links, heterogeneous onboard resources, and time-varying traffic loads make collaborative task offloading difficult for static or reactive strategies. Although multi-agent reinforcement learning (MARL) provides an adaptive solution, existing model-free MARL methods often suffer from slow convergence, insufficient foresight, and limited robustness in dynamic satellite environments. To address these challenges, this paper proposes a World Model-Based Multi-Agent Proximal Policy Optimization (WM-MAPPO) framework for space computing power networks. The offloading problem is formulated as a partially observable multi-agent decision-making process, where LEO satellites make decentralized decisions under incomplete local observations. A predictive world model learns latent transition dynamics of network states and provides future context for proactive planning. Meanwhile, a Transformer-based policy architecture captures inter-agent dependencies and supports cooperative scheduling under centralized training and decentralized execution (CTDE). Simulation results show that WM-MAPPO achieves higher task completion ratios, lower average latency, improved energy efficiency, and stronger robustness than model-free MARL baselines, heuristic methods, Lyapunov-based scheduling, and MINLP-inspired optimization.
Yuqi Cong, Zhiwei Wei, Jiarui Chen et al.· IEEE Transactions on Cogniti...· 0 citations
A Prioritized Adaptive Weighting based on Deep Deterministic Policy Gradient (PAW-DDPG) as an enhanced Deep Deterministic Policy Gradient (DDPG) algorithm to minimize both processing delay and energy consumption by jointly optimizing user scheduling, partial-task offloading, and UAV trajectory is proposed.
W. Saber, Hanan Algamil, Fifi Farouk et al.· Future Internet· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.