Skip to content

Adaptive Undulatory Locomotion of Snake-like Robots in Dynamic Viscous Environments via Deep Reinforcement Learning

Jul 2026 · arXiv.org · Vol abs/2607.21960 · 0 citations · 35 references
Computer Science

Abstract

This paper demonstrates how deep reinforcement learning (DRL) enables adaptive locomotion of snake-like robots in dynamically changing viscous environments, overcoming the inherent performance limitations of classical predefined control methods. The lack of direct onboard sensors for fluid properties necessitates formulating this task as a partially observable Markov decision process. By employing an asymmetric actor-critic framework, a teacher policy trained using privileged information available only in the physics simulator distills its knowledge into a student policy that relies solely on proprioceptive sensor information. Simulation results across a wide range of dynamic viscosity changes ($10^{-7}$ to $10^{-2} m^2/s$) reveal that the DRL agent autonomously acquires non-sinusoidal adaptive gaits. These gaits improve propulsion velocity and transport efficiency, breaking the inherent limits of conventional sinusoidal and kinematic control. The findings establish that implicit environment inference via privileged information distillation is an effective approach to bypass the constraints of classical models under unpredictable fluid dynamics.

View source

Similar papers

Open access Jul 2026

From insect behavior to transferable robot locomotion: inferring embodied locomotor principles from limited data via adversarial inverse reinforcement learning

Insect locomotion exhibits remarkable adaptability and flexibility despite the limited scale of its nervous system. However, the underlying principles that govern leg coordination remain difficult to extract and model computationally. Understanding how insects achieve stable and adaptive locomotion has long provided important inspiration for the development of control strategies in bio-inspired robotics. Nevertheless, many existing approaches rely on predefined coordination rules, manually tuned parameters, or hand-crafted reward functions, which limit the flexibility and transferability of the resulting control strategies. To address this limitation, this study proposes a data-driven framework based on adversarial inverse reinforcement learning , which directly learns continuous locomotion control policies from stick insect walking data and infers latent reward structures and control strategies from biological behavioral demonstrations. Experimental results show that even when trained using only a short segment of flat-terrain demonstration data, the learned policy can still be extracted to learn adaptive and flexible leg coordination patterns under different environmental conditions. Furthermore, the learned reward network can be transferred across different dynamic systems to guide policy learning for robot models with different morphologies. Compared with methods relying solely on reward shaping, the proposed approach achieves faster convergence and produces more biologically consistent gait coordination. A preliminary deployment on a physical bio-inspired robot further demonstrates the potential of the learned policy for sim-to-real application. The proposed method provides a transferable data-driven framework for extracting and learning locomotion control strategies from biological behavior, with potential applications to bio-inspired robotic systems.

Yuchen Wang, Thirawat Chuthong, M. Hayashibe et al. · 0 citations
Sep 2026

Extending the Speed Limit of Quadrupedal Locomotion via Refined Actuator Modeling and Adaptive Command Scheduling

Achieving high-speed locomotion in quadrupedal robots remains highly challenging, as actuators operate near their physical limits and exhibit pronounced nonlinearities. However, many existing methods neglect actuator nonlinearities and physical constraints during training, leading to a significant sim-to-real gap under highly dynamic motions and limiting achievable performance. To address this issue, we propose a high-speed locomotion framework that reduces sim-to-real discrepancies and stabilizes learning over a wide command distribution. A refined actuator model explicitly captures high-speed voltage coupling and magnetic saturation, enabling a more accurate representation of the torque-speed envelope. In addition, a reinforcement learning framework incorporating a two-stage curriculum and adaptive command scheduling (ACS) ensures stable training. Experiments on the 36.5 kg quadruped BlackPanther2 (BP2) demonstrate speeds of up to 13.2 m/s on a treadmill and 11.65 m/s outdoors, establishing a new state-of-the-art and, to the best of our knowledge, a world record for quadrupedal robot locomotion. The results further highlight the importance of accurate actuator modeling in preventing non-physical policy exploitation, and show that ACS improves robustness without sacrificing performance.

Yu-Cheng Tao, Shao-wen Cheng, Guo-Rong Lan et al. · 0 citations
Aug 2026

A computational framework for Kármán gaiting in robotic fish: spatio-temporal perception and CPG-based reinforcement learning

A fully computational framework focusing on the modeling and simulation of a spatio-temporal sensory system to autonomously generate the Kármán gait is proposed, providing a robust algorithmic blueprint for future physical deployments in complex aquatic environments.

Xin-Qi Wang, Ming Wang, Xin-Yan Liu et al. · 0 citations
Conference Aug 2026

End-to-End Control of a Quadruped Robot Using Deep Reinforcement Learning

This paper presents the development and implementation of an end-to-end control framework for a quadruped walking robot based on deep reinforcement learning. The primary objective of the study is to design and verify a control system capable of autonomously generating locomotion strategies. A model of the walking robot was developed using the Simscape Multibody toolbox, providing a physics-based simulation environment for training and evaluation. The proposed control approach employs a deep reinforcement learning agent trained using the Twin Delayed Deep Deterministic Policy Gradient (TD3) algorithm. The agent learns locomotion behaviors directly from interactions with the simulated environment, without relying on predefined gait trajectories or manually designed control laws. Through iterative training, the agent optimizes its policy to maximize a predefined reward function, enabling the robot to discover efficient and stable movement patterns. Simulation results demonstrate that the TD3-based approach is highly effective for continuous control tasks involving systems with complex nonlinear dynamics. The trained agent successfully learned locomotion strategies, including dynamic gaits with flight phases, highlighting the ability of reinforcement learning methods to handle naturally unstable behaviors that are difficult to design using classical control techniques.

Filip Połatyński, Paweł Skruch · 0 citations
Open access Aug 2026

Application and evaluation of reinforcement learning for two-dimensional trajectory tracking in snake-like robots

Context—Snake-like robots are biomimetic systems that can move effectively in narrow, complex, and restricted environments thanks to their modular and flexible body structures composed of numerous serially connected joints. These characteristics offer significant advantages, particularly in areas such as pipeline inspection, search and rescue operations, industrial maintenance applications, and exploration missions. The multiple degrees of freedom distributed along the body enable the robot to achieve high maneuverability but also make the control problem quite complex. Due to the dynamic interactions between segments, friction-based motion characteristics, and nonlinear system behavior, achieving reliable and accurate trajectory tracking emerges as a significant engineering problem.Objective—In this study, a reinforcement learning (RL) based control method has been developed to solve the trajectory tracking problem for snake-like robots in a two-dimensional plane.Method—In the proposed approach, the robot’s dynamic model was created in the Webots simulation environment, an open-source simulation program, and all training and testing processes were carried out in this environment. During the learning process, policy- based RL algorithms from the Stable-Baselines library were used. In this context, Proximal Policy Optimization (PPO) and three different RL algorithms were used during the training process. To enable the robot to adapt to different orientation scenarios, seven different angles defined in the range of +45 to −45 and trajectories of varying lengths were used. Thus, the goal was for the agent to learn a generalizable control policy not only for a specific trajectory type but also for tracks with different slopes and orientations.Results—The results obtained show that the PPO algorithm produced a higher average reward compared to other methods and exhibited a more stable learning process. After training was completed, the developed method was tested both on trajectories used during the training phase and on previously unseen trajectories. For the 0 trajectory, maximum errors were recorded as 0.093 m and 0.040 m for the x and y axes, respectively. Furthermore, the system exhibited robust generalization capabilities on a +22.5 trajectory, not encountered during the training phase, yielding maximum errors of 0.099 m and 0.052 m.Conclusion—These findings demonstrate that the proposed RL-based control approach can effectively solve the two-dimensional trajectory tracking problem in snake robots. In future studies, the proposed method can be extended to the three-dimensional trajectory tracking problem, or it can be evaluated under more complex conditions, such as scenarios involving obstacles.

Furkan Mezgil, M. Bingöl · 0 citations
Preprint Aug 2026

Learning Fault-Tolerant Locomotion with Adaptive Gait Timing

Hardware failures require legged robots to rapidly reorganize coordination and gait timing to maintain stability and mobility. This is particularly challenging for larger quadrupeds, where increased mass and tighter actuation limits reduce the feasibility of aggressive, high-frequency compensation strategies often observed on smaller platforms. In this work, we propose a deep reinforcement learning approach for fault-tolerant locomotion under actuator power loss. The method employs an asymmetric actor-critic architecture in which the critic has access to privileged information during training, while the actor learns to reconstruct a corresponding latent representation from proprioceptive observations. We introduce a latent-alignment loss that encourages consistency between actor and critic representations. Additionally, we augment the action space with a learnable gait frequency parameter, enabling adaptive gait timing in response to terrain variations and actuator degradation without predefined faulty-leg strategies. The approach is validated in high-fidelity simulation on uneven terrain and real-world experiments on flat ground using a 68 kg quadruped robot.

Giovanbattista Gravina, Luca Rossini, Carlo Rizzardo et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.