SR-TD3: Safe-Trajectory-Prior Recurrent TD3 for UAV-Aided Relay Trajectory Design under Dynamic Jamming
Abstract
Uncrewed aerial vehicle (UAV) relays can extend low-altitude coverage, but dynamic jammers make reliable trajectory planning difficult. Deep reinforcement learning (DRL) is widely applied to path planning problems, but existing methods still face key challenges in efficient and risk-aware path learning for observation-constrained agents in dynamic scenarios. We formulate the task as a continuous-action Markov decision process (MDP) and propose safe-trajectory-prior recurrent TD3 (SR-TD3). The proposed SR-TD3 method combines temporal state modeling, separated reward-cost learning, hazard-aware constrained optimization, and training-only safe-trajectory prior regularization. This design links temporal degradation cues, constraint learning, and self-generated behavior priors without adding a separate deployment controller. Simulations under three jammers show that SR-TD3 achieves a global-average reward of 307.80 and a success rate of 81.30%, outperforming TD3 by 64.4% and 55.1%, respectively. The results indicate that the proposed method can effectively improve empirical relay reliability and anti-jamming robustness in dynamic interference environments.