Sim-to-Real Yaw Control of a Robotic Sea Lion Using Deep Reinforcement Learning
Abstract
Yaw regulation of biomimetic underwater robots is complicated by flexible body motion, nonlinear hydrodynamics, and coupled actuation. This study examines whether a policy trained in simulation can be deployed on an existing robotic sea lion (RSL) without changing its hardware or low-level controllers. A deep deterministic policy gradient (DDPG) controller was formulated from measurable states and available actuator commands and trained in Webots using a model calibrated from previous tank tests. Four manually selected reward-weight settings and command update rates of 1, 2, 5, and 10 Hz were examined as deployment-oriented sensitivity comparisons, after which a 5 Hz policy was evaluated in six tank trials. Performance was reanalyzed using the circular-angle mean absolute error (MAE) and root mean square error (RMSE). In the two straight-swimming trials, the per-trial MAE was 1.01–1.12°, and the RMSE was 1.10–1.46°. In the four turning trials, evaluated from the first target crossing to the end of each record, the MAE was 1.71–3.63°, the RMSE was 2.10–4.20°, and the maximum overshoot was 2.99–7.70°. Despite the transient differences between simulations and experiments, the controller regulated the robot toward the target headings in all six tank trials. These results demonstrate the successful sim-to-real deployment of reinforcement-learning-based yaw control under low-frequency communication constraints and provide experimental evidence for its application to biomimetic underwater robots.