A DDPG-Based Offline–Online Joint Optimization Framework for RIS-Assisted UAV-Enabled ISEC in Dynamic Vehicular Networks
Abstract
Intelligent transportation systems face dual challenges in achieving efficient synergy between sensing and edge computing, arising from dynamic environments and unstable communication channels. Unlike existing approaches that rely on perfect channel state information (CSI) or assume static environments, this paper proposes a deep deterministic policy gradient (DDPG)-based offline-online joint optimization (OOJO) framework for trajectory optimization and resource allocation in a reconfigurable intelligent surface-assisted unmanned aerial vehicle (UAV) integrated sensing and edge computation system. Imperfect CSI and highly dynamic vehicle mobility are explicitly modeled, formulating a mixed-integer non-convex problem to minimize the total sum of the maximum end-to-end latency across all UAVs. The proposed centralized single-agent OOJO framework employs an actor-critic network to output continuous control actions, enabling closed-loop optimization from sensing to scheduling. During offline learning, general policies are learned from historical data; during online learning, strategies are rapidly adapted to dynamic environmental changes. Simulation results show that the proposed DDPG algorithm outperforms the Deep Q-Network by 21.2% while maintaining low latency across varying vehicle speeds and imperfect CSI, validating its practicality and efficiency in complex dynamic traffic scenarios.