Skip to content
Conference

End-to-End Control of a Quadruped Robot Using Deep Reinforcement Learning

Aug 2026 · International Conference on Methods & Models in Automation & Robotics · pp. 303-308 · 0 citations · 10 references

Abstract

This paper presents the development and implementation of an end-to-end control framework for a quadruped walking robot based on deep reinforcement learning. The primary objective of the study is to design and verify a control system capable of autonomously generating locomotion strategies. A model of the walking robot was developed using the Simscape Multibody toolbox, providing a physics-based simulation environment for training and evaluation. The proposed control approach employs a deep reinforcement learning agent trained using the Twin Delayed Deep Deterministic Policy Gradient (TD3) algorithm. The agent learns locomotion behaviors directly from interactions with the simulated environment, without relying on predefined gait trajectories or manually designed control laws. Through iterative training, the agent optimizes its policy to maximize a predefined reward function, enabling the robot to discover efficient and stable movement patterns. Simulation results demonstrate that the TD3-based approach is highly effective for continuous control tasks involving systems with complex nonlinear dynamics. The trained agent successfully learned locomotion strategies, including dynamic gaits with flight phases, highlighting the ability of reinforcement learning methods to handle naturally unstable behaviors that are difficult to design using classical control techniques.

View source

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.