A Large Language Model-guided cooperative decision-making framework for joint velocity control and bidirectional channel selection in an AAM system with Aerial Vehicles communicating with ground Base Stations while following predefined linear routes is proposed.
Abstract
In Advanced Air Mobility (AAM) applications, jointly optimizing multi-agent motion control along predefined flight routes and spectrum access is highly challenging due to the tight coupling among mobility, interference, and safety constraints under limited spectrum resources. This paper proposes a Large Language Model (LLM)-guided cooperative decision-making framework for joint velocity control and bidirectional channel selection in an AAM system with Aerial Vehicles (AVs) communicating with ground Base Stations (BSs) while following predefined linear routes. We formulate the problem as a cooperative Markov game with a discrete action space that includes both velocity selection and uplink and downlink channel access, while satisfying Signal-to-Interference-plus-Noise Ratio (SINR) quality requirements and collision avoidance constraints. To obtain reliable expert behavior, we first learn a near-optimal policy using Multi-Agent Reinforcement Learning (MARL) with Value Decomposition Dueling Double Deep Q-Networks (VD3QN). We then treat joint decision-making as a sequence generation task and employ Large Language Models (LLMs) to generate complete joint action sequences from structured environment descriptions, under both Prompt Engineering (PE) and Parameter-Efficient Fine-Tuning (PEFT) via Low-Rank Adaptation (LoRA) on expert demonstrations. Extensive simulations show that structured prompting improves decision quality, while LoRA fine-tuning further increases reward, reduces variance, and yields decisions that closely match the expert policy. Beyond this imitation role, the LLM layer turns 6-AV expert demonstrations into a sequence-level decision generator. In an unseen 10-AV scenario, this generator achieves stronger zero-shot generalization than the VD3QN policy transferred from the 6-AV environment.
Unmanned aerial vehicle (UAV)-assisted mobile edge computing (MEC) systems provide flexible computing services for resource-constrained devices, but malicious jamming attacks introduce dynamic channel conditions and resource competition, making joint trajectory and resource optimization challenging. This paper investigates this problem in multi-UAV MEC systems under jamming, aiming to minimize delay and energy consumption while ensuring anti-jamming robustness. The problem is formulated as a decentralized partially observable Markov decision process (Dec-POMDP). However, traditional multi-agent reinforcement learning (MARL) approaches struggle with high exploration costs and low sampling efficiency in high-dimensional hybrid action spaces. To overcome these limitations, we propose an LLM-guided MARL framework instantiated with the multi-agent deep deterministic policy gradient (MADDPG), which leverages LLM-generated semantic trajectory prompts to dynamically constrain exploration within the continuous action space, effectively compressing the policy search space and accelerating convergence. Simulation results demonstrate that the proposed method achieves $3.4\times $ to $5\times $ faster convergence over hierarchical MADDPG, MADDPG, and independent soft actor-critic (ISAC) baselines, significantly reducing training costs while maintaining superior performance and anti-jamming robustness.
Yeguang Qin, Jie Tang, Fengxiao Tang et al.· IEEE Transactions on Communi...· 0 citations
This study investigates the optimization of three-dimensional (3D) trajectory planning and resource allocation in unmanned aerial vehicle (UAV)-enabled wireless networks with no-fly zones (NFZs) using a deep learning framework. The objective is to maximize the minimum average spectral efficiency (SE) among mobile users served by multiple UAVs while addressing key challenges, including interference from concurrent UAV transmissions, collision avoidance, and NFZ constraints. A realistic probabilistic channel model is considered, where the likelihood of a line-of-sight (LoS) condition is modeled as a function of the elevation angle in the air-to-ground (A2G) link. To solve the formulated optimization problem, a novel deep learning framework with specialized deep neural network (DNN) structures is developed. This framework jointly optimizes 3D UAV trajectory planning and resource allocation, employing an unsupervised learning-based training approach that eliminates the need for labeled data. Performance evaluations demonstrate that the proposed scheme effectively accounts for the probabilistic channel model and co-channel interference while accounting for collision avoidance and NFZ-related constraints. Moreover, it outperforms baseline methods by achieving a higher minimum average SE with real-time computational efficiency, making it practical for UAV-assisted wireless networks.
The proposed multi-agent reinforcement learning policy attains slightly higher throughput with fewer handovers by offloading a fraction of the users to the MEO and GEO layers, an emergent multi-orbit behavior that drives its favorable throughput and handover trade-off.
Yassine Afif, Ashutosh Balakrishnan, Philippe Martins et al.· 0 citations
Unmanned Aerial Vehicles (UAVs) are promising relay platforms due to their flexible deployment and high probability of line-of-sight (LoS) connectivity. This paper compares three deep reinforcement learning (DRL) algorithms-Proximal Policy Optimization (PPO), Soft Actor-Critic (SAC), and Recurrent PPO with LSTM memory-for joint UAV trajectory and energy optimization in UAV based relay systems. The problem formulated is a non-convex optimization problem that minimizes UAV propulsion energy while satisfying Quality of Service (QoS) and mobility constraints under realistic 3GPP channel conditions. Simulation results show that all methods achieve over 99% QoS satisfaction. SAC exhibits the fastest convergence, whereas the proposed Recurrent PPO achieves the lowest energy consumption (44.72 kJ), reducing energy usage by 5.1% compared with PPO. These results highlight the trade-off between convergence speed and energy efficiency in DRL-based UAV relay optimization.
Aniket Subbanwar, Ojas Joshi, Amit Agarwal· International Conference on...· 0 citations
This paper proposes a heterogeneous multi-agent proximal policy optimization (MAPPO)-based framework where both user devices and UAVs act as heterogeneous agents and utilizes a centralized training and decentralized execution (CTDE) paradigm to enable collaborative strategies between computing requesters and providers.
Ming Cheng, Canlin Zhu, Jiang-Hang Tang et al.· Journal of King Saud Univers...· 0 citations
Unmanned Aerial Vehicles (UAVs) are pivotal for facilitating data collection in emergency scenarios. Despite the potential of Multi-Agent Deep Reinforcement Learning (MADRL) in coordinating such systems, existing researches struggle to resolve the high-dimensional coupling of data collection, trajectory planning, and energy scheduling under strict collision avoidance and Return-To-Base (RTB) constraints. This paper proposes a energy-aware cooperative MADRL framework designed to maximize data collection utility under energy constraints. Specifically, we employ a Multi-Agent Deep Deterministic Policy Gradient (MADDPG) approach featuring a Centralized Training with Decentralized Execution (CTDE) design and a multi-objective reward mechanism to balance conflicting optimization goals. Extensive simulations validate the advantages of the proposed framework over leading baselines. Notably, the algorithm exhibits significant quantitative advantages in complex high-load scenarios. These outcomes prove that our method achieves higher task completion rates while strictly adhering to RTB and safety protocols.
Jing Mei, Jing-Lei Xu, Zhao Tong et al.· IEEE Transactions on Network...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.