Nov 2025· IEEE Transactions on Green Communications and Networking· Vol 10, pp. 3742-3755· 0 citations· 31 references
Computer ScienceEngineering
TL;DR
Experimental results show the spatiotemporal A2C policy outperforms IA2C, ConseNet, FPrint, DIAL, and CommNet, achieving faster convergence, higher asymptotic reward, reduced temporal-difference(TD), advantage estimation errors, and a better communication throughput–energy trade-off.
Abstract
In post-disaster space–air–ground integrated networks (SAGINs), terrestrial infrastructure is often impaired, and unmanned aerial vehicles (UAVs) must rapidly restore connectivity for mission-critical ground terminals in cluttered non-line-of-sight (NLoS) urban environments. To enhance coverage, UAVs employ movable antennas (MAs), while reconfigurable intelligent surfaces (RISs) on surviving high-rise buildings redirect signals. The key challenge is communication-limited partial observability, leaving each UAV with a narrow, fast-changing neighborhood view that destabilizes value estimation. Existing multi-agent reinforcement learning (MARL) approaches are inadequate, in that non-communication methods rely on unavailable global critics, heuristic sharing is brittle and redundant, and learnable protocols (e.g., CommNet, DIAL) lose per-neighbor structure and aggravate non-stationarity under tight bandwidth. To address partial observability, we propose a spatiotemporal A2C where each UAV transmits prior-decision messages with local state, a compact policy fingerprint, and a recurrent belief, encoded per neighbor and concatenated. A spatial discount shapes value targets to emphasize local interactions, while analysis under one-hop-per-slot latency explains stable training with delayed views. Experimental results show our policy outperforms IA2C, ConseNet, FPrint, DIAL, and CommNet, achieving faster convergence, higher asymptotic reward, reduced temporal-difference(TD), advantage estimation errors, and a better communication throughput–energy trade-off.
Space-air-ground integrated networks (SAGINs) can provide ubiquitous and reliable connectivity for unmanned aerial vehicles (UAVs). However, air-to-ground links, which are typically dominated by line-of-sight (LoS) propagation, are vulnerable to passive eavesdropping due to the broadcast nature of wireless channels. To enhance physical-layer security, we investigate a SAGIN-enabled secure downlink communication system in which UAVs select service links among satellite, aerial, and terrestrial networks while adjusting the positions of the movable antenna (MA) array to fully exploit connectivity and spatial degrees of freedom for improved secrecy communication performance. Specifically, we maximize the secrecy energy efficiency (SEE) of a UAV swarm by jointly optimizing the MA positions, UAV trajectories, and link selections, subject to UAV mobility, MA movement, and link connectivity constraints. To reduce the real-time channel state information (CSI) acquisition overhead, we propose a channel knowledge map (CKM)-assisted multi-agent reinforcement learning framework. Specifically, the CKM is first constructed from sparse channel measurements via Kriging interpolation and is then leveraged together with satellite ephemeris information to enable efficient storage and retrieval of CSI. To reduce the action-space dimensionality and computational complexity, we model the MA array using rigid-body kinematics and adjust its position through global rigid-body translation, thereby constructing a low-dimensional hybrid action space for the joint optimization decisions. To align local decisions with system-wide performance under system constraints, we design an individual-team collaborative reward mechanism and introduce action masks to enforce constraints on UAV mobility, collision avoidance, MA regions, and connectivity capacity.
Jiayang Wan, Ya-Fei Wang, Jiawei Zhuang et al.· 0 citations
Simulation results demonstrate that RESCUE-ISAC improves energy efficiency, link reliability, sensing performance, mobility robustness, and runtime–performance trade-off compared with heuristic, lightweight, and optimization-based benchmark schemes.
R. Khalil, Saba Mahmood, T. Jan et al.· IEEE Open Journal of Vehicul...· 0 citations
Sustaining service for survivor IoT devices in disaster zones where terrestrial infrastructure has failed is an open challenge in sixth-generation (6G) emergency communications. This paper studies a single UAV flying base station that serves survivor devices over power-domain NOMA, with a passive reconfigurable intelligent surface (RIS) providing a reflected bypass for far survivors blocked by deep rubble obstruction, and a Deep Q-Learning (DQL) agent placing the UAV in 3D. We establish three results. First, the conventional sum-rate positioning objective is structurally blind to blocked survivors: because the far users contribute only a negligible fraction of the optimised aggregate throughput, the positioning reward is dominated by the near users, so a throughput-optimal controller abandons the very survivors the network exists to reach. Second, we expose the governing geometry: for a blocked survivor the achievable rate is dictated by the UAV-to-RIS separation rather than the UAV-to-survivor separation (because only the UAV-to-RIS hop of the two-hop cascade is controllable), yielding a counterintuitive design rule: to serve a blocked survivor, move the UAV toward the reflector, not the survivor. The throughput-optimal and survivor-optimal UAV positions are consequently distinct, creating an irreducible positioning tension. Third, because a single passive RIS cannot phase-align to all blocked survivors simultaneously, we introduce an on-demand, per-beacon RIS scheduling model and characterise its delivery capacity with a queuing analysis, resolving a physical contradiction common to multiuser RIS studies. We cast the throughput-reliability tension as a Pareto frontier and resolve it with a survivor-aware reward; re-trained under this reward, the DQL agent autonomously reaches a position delivering near-100% far-user telemetry-beacon reliability at an aggregate-throughput cost of about 7%. All findings are reproduced by openly available simulation code. The study is positioned as a sub-6 GHz emergency-connectivity framework for future 6G architectures: the rates reported here are specific to a 3.5 GHz, 20 MHz disaster-coverage configuration and are not intended to represent general 6G performance.
G. Dayasagar, S. Nivethika· IEEE Access· 0 citations
This paper proposes AERIS, an offline policy improvement framework for multi-UAV ISAC that learns from fixed flight logs under centralized training and decentralized execution and designs STAR-CRDT, an offline multi-agent RL algorithm that performs support-aware local action rectification and distills only trusted improvements into the decentralized actor.
Ziyuan Wang, Yi-Fan Sui, Wei Wei et al.· 0 citations
In disaster response and other infrastructure-limited settings, UAV-mounted access points can rapidly restore service availability for mobile ground users as demand and fleet availability evolve. Existing single-slot coverage formulations, however, can mask prolonged individual outages and do not jointly represent heterogeneous service priorities, finite battery capacities, and periodic recharging. We study persistent geometricmulti-UAV service coverage, where a user is available for service when it lies inside a UAV footprint. We propose Priority- and Outage-Guided Safe QMIX (POGS-QMIX), a hybrid hierarchical framework in which a centralized online coordinator forms conflict-reduced UAV–user targets from fleet-wide priority and outage information, while parameter-shared QMIX agents independently choose target-conditioned low-level actions. The framework couples class-balanced outage memory, assignment, dense target-progress feedback, and a return-energy action mask. The evaluation includes learning and non-learning baselines, greedy-versus-Hungarian assignment, multi-seed statistics, sensitivity studies, operating-condition studies, and energy-stress tests. In the default scenario, POGS-QMIX obtains high-priority coverage 0.547±0.009 and maximum high-priority outage 38.0±4.7 slots over five independent seeds.
Hao-Yu Mei, Cheng-Tao Xu, Ruo-Zhe Li et al.· Drones· 0 citations
Rapid, reliable, and energy-efficient data collection is essential for disaster response, where terrestrial communication networks may be disrupted or unavailable. Unmanned Aerial Vehicles (UAVs) provide a flexible means of collecting critical sensing data, but their operation is constrained by limited onboard energy, stochastic wireless conditions, complex three-dimensional environments, and stringent latency requirements. This paper presents a structured multi-UAV framework that separates mission optimisation into spatial, temporal, and safety layers. In the spatial layer, a three-dimensional Travelling Salesman Problem with Neighbourhoods (3D-TSPN) formulation enables UAVs to collect data by entering valid sensing regions rather than visiting exact sensor coordinates. An Age of Information (AoI)-aware Genetic Algorithm (GA) optimises the sensor-visitation sequence, while Rapidly Exploring Random Tree Connect (RRT-Connect) generates obstacle-aware feasible paths in the three-dimensional environment. In the temporal layer, a Lyapunov-based controller selects between local processing and binary offloading to a single Mobile Edge Computing (MEC) node according to queue backlog, processing delay, energy consumption, information freshness, wireless-link feasibility, and task deadlines. In the safety layer, continuous-time conflict detection and bounded temporal or spatial adjustments are used to monitor and mitigate inter-UAV and obstacle-related risks. The framework is evaluated under stochastic wireless, mobility, computation, and obstacle conditions using 20 independent random seeds. Across the corresponding 20 proposed-policy runs, it achieves a 100% mission-validity rate, complete sensor coverage, no dropped tasks, and zero final collision or near-miss events. Compared with planning-oriented and MEC-oriented baselines, the proposed framework achieves lower information age, average delay, processing delay, energy consumption, and system cost under the evaluated conditions, while maintaining reliable multi-UAV coordination. The layered design also clarifies the contribution of each component: 3D-TSPN provides spatial flexibility, the AoI-aware GA improves route sequencing, RRT-Connect supports obstacle-aware path feasibility, Lyapunov control enables queue-aware processing decisions, and safety monitoring supports coordinated multi-UAV operation. These results indicate that integrating spatial planning, computation control, and safety coordination within a clearly separated layered architecture can provide an effective solution for multi-UAV disaster-response data collection in complex three-dimensional environments.
Rakan Armoush, Shidrokh Goudarzi, Muhammad Nadeem Khan et al.· Italian National Conference...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.