Skip to content

Hybrid DRL-Based Sensing and Age of Information Optimization for UAV-Enabled ISCC

2026 · IEEE Communications Letters · Vol 30, pp. 2450-2454 · 0 citations · 18 references
Computer Science

Abstract

In low-altitude economy (LAE), the deployment of unmanned aerial vehicles (UAVs) provides substantial convenience and enhances operational efficiency. This letter investigates a joint resource and trajectory optimization problem in a UAV-enabled integrated sensing, computation, and communication (ISCC) network where the UAV senses and processes the status information from sensing targets (STs), and then sends the computed results to the data collection center (DC). Aiming to maximize the sensing data volume while minimizing the age of information (AoI), we jointly optimize the UAV’s sensing schedule, number of sensing trials, time allocation, transmit power, CPU frequency, and trajectory. The problem is formulated as a Markov decision process (MDP), and a hybrid deep reinforcement learning (DRL) framework is proposed to derive optimal policies. Specifically, we adopt the twin delayed deep deterministic policy gradient (TD3) framework and enhance it with a hybrid-baseline prioritized experience replay (PER) mechanism, denoted as HPTD3. Simulation results demonstrate that the proposed approach significantly outperforms benchmark schemes.

View source

Similar papers

Conference Jul 2026

Reinforcement Learning-Based Decode-and-Forward UAV Relay Trajectory Optimization

Unmanned Aerial Vehicles (UAVs) are promising relay platforms due to their flexible deployment and high probability of line-of-sight (LoS) connectivity. This paper compares three deep reinforcement learning (DRL) algorithms-Proximal Policy Optimization (PPO), Soft Actor-Critic (SAC), and Recurrent PPO with LSTM memory-for joint UAV trajectory and energy optimization in UAV based relay systems. The problem formulated is a non-convex optimization problem that minimizes UAV propulsion energy while satisfying Quality of Service (QoS) and mobility constraints under realistic 3GPP channel conditions. Simulation results show that all methods achieve over 99% QoS satisfaction. SAC exhibits the fastest convergence, whereas the proposed Recurrent PPO achieves the lowest energy consumption (44.72 kJ), reducing energy usage by 5.1% compared with PPO. These results highlight the trade-off between convergence speed and energy efficiency in DRL-based UAV relay optimization.

Aniket Subbanwar, Ojas Joshi, Amit Agarwal · 0 citations
Jul 2026

CRB-Driven Beamforming and Trajectory Optimization for UAV-assisted ISAC System

Simulation results demonstrate that the proposed method significantly reduces the time-averaged CRB by over 10%, compared with the ISAC system without UAV assistance, and also achieves a higher sensing accuracy than both the fixed-UAV-trajectory and the maximum-ratio-transmission-based beamforming benchmarks.

Yi Yang, Qianqian Zhang, Huaxia Wang · 0 citations
Jul 2026

Reinforcement Learning-Driven Optimal Uav Selection Framework for Efficient Uav-To-Uav Communication

Unmanned Aerial Vehicles (UAVs) have gained widespread attention in diverse applications like military, medical, aerial surveillance and many more. Presently, the problem of limited bandwidth and geographic factors has raised the need for effective and timely data transfer. Training UAVs with reinforcement learning-based algorithms facilitates autonomous decision-making capabilities. In this paper, we proposed an intelligent system for the optimal UAV selection process by evaluating the continuous performance of each UAV. The analyzing factors are based on the real-world factors affecting the quality of signals, such as noise interference, relative motion between source and wave, and transmission power. Based on the systematic conditions observed, the system provides efficient rewards. To promote the selection of the optimal UAV and enhance the learning process, the state information of the UAV is fed into a deep neural network (DQN), which predicts the 'Q-values'. Our system implements a deep Q-learning algorithm, which enhances the agent's performance by systematically learning from its experience. The model operates accurately by selecting the most reliable UAV, thus, enhancing the throughput by optimal power allocation. It outperforms other conventional models in terms of timely data delivery and energy utilization. The system adapts various complex patterns by analyzing the historical and present scenarios. Empowered by this intelligent system, time-critical decision-making can be achieved with minimal energy consumption.

Divyanshu Bhardwaj, Angel Kanjiya, N. Jadav et al. · 0 citations
Preprint Aug 2026

AERIS: Offline Policy Improvement for Multi-UAV Integrated Sensing and Communication

This paper proposes AERIS, an offline policy improvement framework for multi-UAV ISAC that learns from fixed flight logs under centralized training and decentralized execution and designs STAR-CRDT, an offline multi-agent RL algorithm that performs support-aware local action rectification and distills only trusted improvements into the decentralized actor.

Ziyuan Wang, Yi-Fan Sui, Wei Wei et al. · 0 citations
Open access 2026

Fast Learning for Optimization of Green Edge Collaborative UAV

A UAV motion-aware image capturing and communication system that dynamically optimizes data offloading by jointly considering scenario variability and communication resource allocation and a fast learning-based optimization algorithm (FLO-MICC).

Xiaoqing Liu, Songtao Gao, Qixuan Zhang et al. · 0 citations
Jul 2026

UAV Swarming for Air-Ground ISAC via Cross-Region Cooperation

A service-driven regional partitioning scheme is proposed to support traffic-aware UAV communication, and an adaptive handshaking mechanism is introduced to improve cooperative sensing accuracy by mitigating residual inter-region phase errors with controlled synchronization overhead.

Linghui Miao, Shijian Gao · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.