AUVs Path Planning Based on DSAC-T With Reward-Adaptive PER
Efficient 3-D path planning for autonomous underwater vehicles (AUVs) in dynamic submarine environments presents a significant challenge due to complex seabed terrain, ocean currents, and obstacles. In view of the adaptability and generalization limitations of traditional methods, this article proposes the reward-adaptive prioritized experience replay (RAPER) mechanism to dynamically adjust the experience sampling priorities by evaluating the temporal-difference errors and the weights of critical reward events, including success rate, goal proximity, route length, obstacle avoidance, and ocean current utilization. Then, a deep reinforcement learning framework, the DSAC-T-RAPER algorithm, is established for underactuated AUVs by combining the RAPER mechanism and distributional soft actor–critic with three refinements (DSAC-T) algorithm. Meanwhile, the finite-step evolution and boundedness of the adaptive event weight are theoretically analyzed to guarantee the reliability and stability of the proposed algorithm. The 3-D simulation environment is constructed by considering complicated submarine seafloor topography, real ocean current data from the Copernicus Marine Environment Monitoring Service, and randomly distributed obstacles. Simulations under different scenarios demonstrate that, compared with some other algorithms, DSAC-T-RAPER achieves higher success rates and better performances.