Skip to content

Reinforcement Learning for Delivery Drone-Based Participatory Sensing in Dynamic Environments

Jul 2026 · arXiv.org · Vol abs/2607.18874 · 0 citations · 37 references
Computer Science

TL;DR

A Two TimeScale Reinforcement Learning framework (TSRL), which separates decision-making into two cooperative layers and significantly outperforms baselines, achieving average system profit improvements of 20.1% in Hangzhou and 46.6% in Shanghai.

Abstract

Using Unmanned Aerial Vehicle (UAV) for urban sensing has emerged as a powerful paradigm to monitor the status of the city, e.g., air quality and noise levels, through agile aerial crowdsourcing. Despite this potential, existing UAV-based sensing approaches overlook environmental disturbances like wind that drastically impact drone velocity and energy efficiency. Consequently, directly applying existing methods to this joint delivery and sensing paradigm in dynamic environments faces two severe challenges: (1) scalability bottlenecks as fleet sizes expand; and (2) multi-timescale decision heterogeneity between macro task dispatching and micro velocity control. To tackle these, we formalize the problem as SensUAV and propose a Two TimeScale Reinforcement Learning framework (TSRL). Specifically, TSRL separates decision-making into two cooperative layers. At the macro level, a task-embedding sensing dispatcher handles scalability by separately encoding distinct task features and sequentially evaluating UAV suitability before task selection. At the micro level, a wind-aware velocity controller learns fine-grained velocity scheduling to adapt to dynamic environmental variations. Extensive experiments on real-world datasets demonstrate that TSRL significantly outperforms baselines, achieving average system profit improvements of 20.1% in Hangzhou and 46.6% in Shanghai.

View source

Similar papers

Preprint Aug 2026

AERIS: Offline Policy Improvement for Multi-UAV Integrated Sensing and Communication

This paper proposes AERIS, an offline policy improvement framework for multi-UAV ISAC that learns from fixed flight logs under centralized training and decentralized execution and designs STAR-CRDT, an offline multi-agent RL algorithm that performs support-aware local action rectification and distills only trusted improvements into the decentralized actor.

Ziyuan Wang, Yi-Fan Sui, Wei Wei et al. · 0 citations

Disturbance-Aware Hybrid Learning for Robust and Adaptive UAV Flight in Extreme Winds

WA-TD3 is introduced, a data-driven control framework that enables real-time wind disturbance perception and adaptive compensation without dedicated wind sensors, and consistently outperforms state-of-the-art methods on tracking accuracy under strong winds.

Hui-Dong Liu, Jiarui Dou, Jiangshan Ai et al. · 0 citations
Open access Aug 2026

SkyAgent: A lightweight LLM-driven reinforcement learning framework for adaptive cooperative path planning of two UAVs

This work provides a feasible technical pathway and reproducible evaluation benchmark for the collaborative deployment of lightweight LLM planner, sub-goal guidance, sensor observations, cooperative reward, and reward shaping components and quantifies the indispensability of the LLM planner.

Yuting Cao, Zheng Zhao, Jiekai Wu et al. · 0 citations
Jul 2026

TRUAV: Distributed Multi-Agent Reinforcement Learning for Trajectory Planning and Routing Enhancement in UAV-Aided IoT-Enabled VANETs

Unmanned aerial vehicles (UAVs) have emerged as a key enabler of next-generation Internet of Things (IoT) ecosystems, offering flexible aerial relaying to extend connectivity across dynamic vehicular ad hoc networks (VANETs) in smart city environments. However, conventional centralized approaches for UAV trajectory planning require continuous global network state aggregation, making them impractical under bandwidth and energy constraints typical of dense urban deployments. In this article, we present TRUAV, a distributed multi-agent reinforcement learning framework based on independent tabular Q-learning for joint UAV trajectory planning and routing enhancement in UAV-aided VANETs. Each UAV is equipped with a local Q-learning agent that operates purely on locally observable information, including vehicle density, packet queue states, and neighbor UAV positions, thereby eliminating the need for global state exchange. A potential-game-inspired reward design encourages spatial diversity and routing-aware UAV positioning among interacting agents while accounting for energy consumption. Numerical simulations over a large urban area with 200 mobile vehicles show that the proposed TRUAV framework achieves network coverage and packet delivery ratios comparable to centralized deep reinforcement learning methods, while also improving relay delay and energy efficiency. Finally, we discuss emerging challenges and future research directions for distributed multi-agent UAV-assisted IoT systems.

Muhammad Umar Farooq Qaisar, Lin Zhang, Zhen Chen et al. · 0 citations
2026

Attention-Enhanced Hierarchical Reinforcement Learning for Air-Ground Cooperative Perception

Air-ground cooperative perception (AGCP) integrates connected and autonomous vehicles (CAVs), roadside units (RSUs), and uncrewed aerial vehicles (UAVs) to provide wide-area coverage and high-resolution perception by leveraging their complementary perception and communication capabilities. However, the dynamic and heterogeneous characteristics of the air-ground network introduce strong cross-layer coupling across perception, communication, and computation, thereby complicating the coordination of cooperation update intervals and cooperation partner selection. To address these challenges, we develop a unified AGCP framework that jointly models LiDAR-based multi-agent perception, together with its associated communication bandwidth allocation and computation latency models, under dynamic mobility and time-varying resource conditions. Building on this framework, a multi-objective optimization problem is formulated to characterize the interplay between update interval selection and cooperation partner choice, aiming to balance perception accuracy and end-to-end latency. A Tchebycheff distance-based formulation is utilized to normalize and integrate multiple objectives into a unified optimization metric. To efficiently solve this highly coupled problem, an attention-enhanced hierarchical reinforcement learning algorithm is proposed, which leverages a two-level Markov decision process combined with an attention-enhanced actor-critic architecture. Simulation results validate that the proposed algorithm achieves a desirable trade-off between perception performance and end-to-end latency.

Hai-Xia Peng, Yixin Fan, Zhou Su et al. · 0 citations
#federated learning Open access Sep 2026

Significance-Aware Federated Reinforcement Learning for AoI Optimization of Vehicular Sensing in UAV-Assisted Edge Networks

Timely vehicular sensing is important for traffic monitoring, cooperative driving, and road-safety management. High mobility, time-varying wireless conditions, and limited edge resources nevertheless make information freshness difficult to maintain. This paper studies age of information (AoI) minimization in a three-layer UAV-assisted edge network comprising vehicle devices (VDs), unmanned aerial vehicles (UAVs), and a cloud center (CC). VDs periodically generate sensor-data packets, UAVs provide mobile edge processing and data-relaying services, and the CC coordinates system-wide resource allocation. The joint optimization of sensor-data transmission, UAV movement, packet processing, computation offloading, and bandwidth allocation is formulated within a cooperative multi-agent framework. To solve this problem, we propose a collaborative heterogeneous federated actor–critic (CHFAC) framework. Its significance-aware federated learning mechanism evaluates local model updates according to update significance, alignment with the global learning direction, and training stability and uses the resulting contribution scores for non-uniform agent selection and contribution-weighted aggregation. In the considered simulation setting, evaluation over 1000 test episodes yields an average AoI of 7.45±1.65 and a worst-case AoI of 38.72±24.06. The average AoI is 79.0%, 63.9%, and 29.2% lower than that obtained by the implemented HF-MARL, H-MAAC, and non-federated baselines, respectively. These results demonstrate the effectiveness of CHFAC for freshness-aware vehicular sensing in dynamic UAV-assisted edge environments.

Xue-Yuan Wang, Si-Yu Bai, Yu Zhang et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.