Skip to content

Optimizing Information Freshness in Satellite-UAV IoRT Networks: A Heterogeneous Multi-Agent Approach

2026 · IEEE Transactions on Cognitive Communications and Networking · Vol 12, pp. 10840-10854 · 0 citations · 52 references

Abstract

In satellite-UAV assisted communication networks, jointly optimizing the UAV’s trajectory and the multi-agent scheduling decisions to minimize the age of information (AoI) is a notoriously challenging problem. The complexity is compounded by the fundamental heterogeneity between the satellite and UAV agents, including their disparate action spaces, partial observations, and differing energy-consumption and communication-cost penalties. To address this, we formulate the problem as a decentralized partially observable Markov decision process (Dec-POMDP) and propose a novel heterogeneous multi-agent compound-action proximal policy optimization (HMACPPO) algorithm. HMACPPO leverages a centralized training with decentralized execution (CTDE) framework, using role-specific decentralized actors together with agent-specific centralized critics conditioned on the global state. Specifically, the UAV employs a compound PPO (CPPO) actor for its hybrid action space, while the satellite uses a PPO actor for discrete scheduling. Extensive simulations show that HMACPPO outperforms the compared baselines, and that the resulting coordinated policy effectively manages the trade-off between AoI, UAV energy consumption, and operational cost.

View source

Similar papers

Open access Jul 2026

A Multi-UAV Cooperative Mission Planning Method Based on Multi-Agent Guided Soft Actor–Critic

Multiple unmanned aerial vehicles (UAVs) performing cooperative missions in complex environments face challenges such as difficult cooperative decision-making, stringent spatiotemporal consistency constraints, and environmental uncertainty. The cooperative mission considered in this paper aims to enable multiple UAVs to simultaneously arrive at multiple constant-velocity moving targets. To address these challenges, this paper proposes a multi-agent guided soft actor–critic (MAGSAC) deep reinforcement learning algorithm. Under the centralized training with decentralized execution (CTDE) framework, a Guider network is introduced to guide the local actor network in learning coordinated strategies, thereby alleviating the non-stationarity of multi-agent decision-making under uncertain environments. An estimated time of arrival (ETA)-based spatiotemporal coordination reward function is designed to promote synchronized arrival. To address sparse rewards, a hindsight experience replay (HER) mechanism based on backward trajectory reconstruction is developed, and a delayed collision-constraint activation mechanism is incorporated to improve convergence while maintaining flight safety. Simulation results show that MAGSAC outperforms existing mainstream algorithms in synchronization success rate, temporal synchronization accuracy, and safety.

Shuanli Jia, Naiming Qi, Zheng Li et al. · 0 citations
Preprint Jul 2026

Multi-Agent Reinforcement Learning for SLA-Aware Network Slicing in UAV-Enabled MEC

A predictive multi-agent Reinforcement Learning (RL) framework that proactively maintains SLA stability in UAV-enabled MEC through coordinated trajectory control and computation resource allocation and designs an SLA-aware reward function that explicitly penalizes both violation probability and duration across slices.

M. Farhoudi, Zeinab Sasan, Masoud Shokrnezhad et al. · 0 citations
Preprint Jul 2026

TRUAV: Distributed Multi-Agent Reinforcement Learning for Trajectory Planning and Routing Enhancement in UAV-Aided IoT-Enabled VANETs

Unmanned aerial vehicles (UAVs) have emerged as a key enabler of next-generation Internet of Things (IoT) ecosystems, offering flexible aerial relaying to extend connectivity across dynamic vehicular ad hoc networks (VANETs) in smart city environments. However, conventional centralized approaches for UAV trajectory planning require continuous global network state aggregation, making them impractical under bandwidth and energy constraints typical of dense urban deployments. In this article, we present TRUAV, a distributed multi-agent reinforcement learning framework based on independent tabular Q-learning for joint UAV trajectory planning and routing enhancement in UAV-aided VANETs. Each UAV is equipped with a local Q-learning agent that operates purely on locally observable information, including vehicle density, packet queue states, and neighbor UAV positions, thereby eliminating the need for global state exchange. A potential-game-inspired reward design encourages spatial diversity and routing-aware UAV positioning among interacting agents while accounting for energy consumption. Numerical simulations over a large urban area with 200 mobile vehicles show that the proposed TRUAV framework achieves network coverage and packet delivery ratios comparable to centralized deep reinforcement learning methods, while also improving relay delay and energy efficiency. Finally, we discuss emerging challenges and future research directions for distributed multi-agent UAV-assisted IoT systems.

Muhammad Umar Farooq Qaisar, Lin Zhang, Zhen Chen et al. · 0 citations
2026

Collaborative Trajectory and Resource Optimization in Multi-UAV MEC Under Jamming: An LLM-Guided MARL Framework

Unmanned aerial vehicle (UAV)-assisted mobile edge computing (MEC) systems provide flexible computing services for resource-constrained devices, but malicious jamming attacks introduce dynamic channel conditions and resource competition, making joint trajectory and resource optimization challenging. This paper investigates this problem in multi-UAV MEC systems under jamming, aiming to minimize delay and energy consumption while ensuring anti-jamming robustness. The problem is formulated as a decentralized partially observable Markov decision process (Dec-POMDP). However, traditional multi-agent reinforcement learning (MARL) approaches struggle with high exploration costs and low sampling efficiency in high-dimensional hybrid action spaces. To overcome these limitations, we propose an LLM-guided MARL framework instantiated with the multi-agent deep deterministic policy gradient (MADDPG), which leverages LLM-generated semantic trajectory prompts to dynamically constrain exploration within the continuous action space, effectively compressing the policy search space and accelerating convergence. Simulation results demonstrate that the proposed method achieves $3.4\times $ to $5\times $ faster convergence over hierarchical MADDPG, MADDPG, and independent soft actor-critic (ISAC) baselines, significantly reducing training costs while maintaining superior performance and anti-jamming robustness.

Yeguang Qin, Jie Tang, Fengxiao Tang et al. · 0 citations
2026

Energy-Aware Multi-UAV Collaboration for Data Collection and Trajectory Planning With MADDPG

Unmanned Aerial Vehicles (UAVs) are pivotal for facilitating data collection in emergency scenarios. Despite the potential of Multi-Agent Deep Reinforcement Learning (MADRL) in coordinating such systems, existing researches struggle to resolve the high-dimensional coupling of data collection, trajectory planning, and energy scheduling under strict collision avoidance and Return-To-Base (RTB) constraints. This paper proposes a energy-aware cooperative MADRL framework designed to maximize data collection utility under energy constraints. Specifically, we employ a Multi-Agent Deep Deterministic Policy Gradient (MADDPG) approach featuring a Centralized Training with Decentralized Execution (CTDE) design and a multi-objective reward mechanism to balance conflicting optimization goals. Extensive simulations validate the advantages of the proposed framework over leading baselines. Notably, the algorithm exhibits significant quantitative advantages in complex high-load scenarios. These outcomes prove that our method achieves higher task completion rates while strictly adhering to RTB and safety protocols.

Jing Mei, Jinglei Xu, Zhao Tong et al. · 0 citations