With the rapid advancement and deep integration of the Internet of Things (IoT) and 5G technologies, mobile edge computing (MEC) has undertaken an increasingly important role in enhancing service quality. Leveraging their high mobility and flexible deployment, unmannedaerial vehicles (UAVs) extend MEC services to challenging environments such as mountainous areas. Nevertheless, UAVs have inherent limitations, including restricted onboard resources (e.g., energy and computing capacity) and the need for obstacle avoidance flight. In this work, which investigates a UAV-assisted MEC system with uneven terrain and dynamic service scenarios, these limitations bring additional challenges to system optimization. The incorporation of terrain information in high-dimensional state space, the continuous action space required for fine control, and the variable network demands under dynamic service scenarios complicate the non-convex optimization problem. By jointly designing UAV’s trajectory and user equipments’ (UEs) task allocation, we address the task offloading problem under safe flight conditions, aiming to maximize both service coverage ratio and UAV’s propulsion energy efficiency. Then, we propose a phased hierarchical deep reinforcement learning (PH-DRL) algorithm, in which the network training is designed in phases and the network structure is organized hierarchically. Specifically, the phased method overcomes insufficient network experience in complex environments, while the hierarchical method decomposes the optimization variables, enabling independent solution. Experimental results demonstrate that the PH-DRL algorithm substantially improves service coverage ratio and propulsion energy efficiency, achieving system utility that significantly outperforms other comparative strategies.
Zhao Tong, Shi-Yan Zhang, Jing Mei et al.· IEEE Transactions on Mobile...· 0 citations
The exponential growth of edge devices and growing demand for low-latency, high-throughput applications have established edge computing as essential infrastructure. Edge computing enables near-data computation to alleviate device burdens, yet task offloading optimization remains challenging due to the heterogeneous, dynamic networks and privacy risks in centralized approaches. To address these intertwined challenges, we propose the federated Q-Learning via Transformer for task offloading (FQTTO) algorithm, which integrates federated learning (FL) with a Transformer-based Q-Network. FL ensures privacy-preserving distributed model updates, while the Transformer employs customized sequential state encodings of task priorities and resource availabilities to autoregressively predict Q-values, incorporating <inline-formula><tex-math notation="LaTeX">$n$</tex-math><alternatives><mml:math><mml:mi>n</mml:mi></mml:math><inline-graphic xlink:href="mei-ieq1-3697303.gif"/></alternatives></inline-formula>-step returns for long-term optimization. Extensive simulations show that compared to the baseline methods, FQTTO achieves average reductions exceeding 20.4% in delay and 22.3% in energy consumption, while improving load balance by at least 13.4% and enhancing the task completion rate by up to 12.1%.
Zhao Tong, Shi-Zhen Xiao, Xi Zhang et al.· IEEE Transactions on Mobile...· 0 citations
Unmanned Aerial Vehicles (UAVs) are pivotal for facilitating data collection in emergency scenarios. Despite the potential of Multi-Agent Deep Reinforcement Learning (MADRL) in coordinating such systems, existing researches struggle to resolve the high-dimensional coupling of data collection, trajectory planning, and energy scheduling under strict collision avoidance and Return-To-Base (RTB) constraints. This paper proposes a energy-aware cooperative MADRL framework designed to maximize data collection utility under energy constraints. Specifically, we employ a Multi-Agent Deep Deterministic Policy Gradient (MADDPG) approach featuring a Centralized Training with Decentralized Execution (CTDE) design and a multi-objective reward mechanism to balance conflicting optimization goals. Extensive simulations validate the advantages of the proposed framework over leading baselines. Notably, the algorithm exhibits significant quantitative advantages in complex high-load scenarios. These outcomes prove that our method achieves higher task completion rates while strictly adhering to RTB and safety protocols.
Jing Mei, Jing-Lei Xu, Zhao Tong et al.· IEEE Transactions on Network...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.