Skip to content

AUVs Path Planning Based on DSAC-T With Reward-Adaptive PER

Sep 2026 · IEEE Systems Journal · Vol 20, pp. 1129-1140 · 0 citations · 32 references

Abstract

Efficient 3-D path planning for autonomous underwater vehicles (AUVs) in dynamic submarine environments presents a significant challenge due to complex seabed terrain, ocean currents, and obstacles. In view of the adaptability and generalization limitations of traditional methods, this article proposes the reward-adaptive prioritized experience replay (RAPER) mechanism to dynamically adjust the experience sampling priorities by evaluating the temporal-difference errors and the weights of critical reward events, including success rate, goal proximity, route length, obstacle avoidance, and ocean current utilization. Then, a deep reinforcement learning framework, the DSAC-T-RAPER algorithm, is established for underactuated AUVs by combining the RAPER mechanism and distributional soft actor–critic with three refinements (DSAC-T) algorithm. Meanwhile, the finite-step evolution and boundedness of the adaptive event weight are theoretically analyzed to guarantee the reliability and stability of the proposed algorithm. The 3-D simulation environment is constructed by considering complicated submarine seafloor topography, real ocean current data from the Copernicus Marine Environment Monitoring Service, and randomly distributed obstacles. Simulations under different scenarios demonstrate that, compared with some other algorithms, DSAC-T-RAPER achieves higher success rates and better performances.

View source

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.