Joint 3D Trajectory and Power Optimization for UAV Swarms in Cell-Free Massive MIMO Networks: A CTDE-MAPPO Framework for Sensing-Aware Precision Agriculture
The integration of Unmanned Aerial Vehicle (UAV) swarms with Cell-Free massive Multiple Input Multiple Output (CF-mMIMO) networks offers promising prospects for large-scale crop monitoring in precision agriculture. CF-mMIMO provides macro-diversity and uniform channel quality across large agricultural fields. However, practical deployment demands jointly optimizing 3D trajectories and transmit power to maximize energy efficiency and field coverage simultaneously. This is challenging due to the limited battery capacity, mandatory return-to-depot constraints, and collision avoidance requirements. In this paper, we introduce a joint sensing–communication utility function that captures the trade-off between energy efficiency and field coverage completeness. To provide a scalable and distributed solution for rotary-wing UAV swarms, we develop a multi-agent deep reinforcement learning (MADRL) methodology based on the Multi-Agent Proximal Policy Optimization (MAPPO) approach. We adopt the Centralized Training with Decentralized Execution (CTDE) strategy, in which a CF-mMIMO central processing unit (CPU) serves as a global critic during training. At execution time, each UAV independently runs a lightweight local policy that adapts its trajectory and transmit power in real time based on battery state and air-to-ground channel variations. Simulation results reveal that the proposed MAPPO-CTDE approach outperforms existing benchmarks. Unlike prior methods that require instantaneous global CSI or neglect the sensing–communication coupling, the proposed approach simultaneously achieves high field coverage completeness, robust communication energy efficiency, and a high depot-return rate under hard battery constraints without any inter-UAV communication overhead at execution time.