Skip to content
Conference

Sim-to-Real Reinforcement Learning for Ball-Balancing Locomotion on Quadruped Robots

Jul 2026 · 2026 IEEE/ASME International Conference on Advanced Intelligent Mechatronics (AIM) · pp. 1-8 · 0 citations · 21 references

Abstract

Non-prehensile manipulation of freely moving objects on a mobile base represents a significant challenge in underactuated robotics. This paper presents a sim-to-real reinforcement learning pipeline for a Unitree Go2 robot tasked with balancing a free-rolling ping-pong ball on a board mounted on its trunk while maintaining stable posture or tracking commanded velocities. Built on Legged Gym and the Genesis simulation, the proposed framework augments standard quadrupedal locomotion with ball-aware observations, task-specific reward design, and curriculum learning for progressively harder balancing and locomotion regimes. To improve transfer, the method incorporates domain randomization, camera-rate-compatible ball observations that mimic asynchronous visual feedback, and deployment-oriented safeguards such as smooth startup action blending. The learned policies are evaluated through a three-stage pipeline: large-scale training in Genesis, sim-to-sim validation in MuJoCo, and deployment on a physical Unitree Go2 using vision-estimated board-frame ball states. Experimental results show that the proposed framework can achieve both standing ball balance and ball-balancing locomotion on hardware, while additional comparisons between PPO and SAC highlight a trade-off between nominal task performance and disturbance robustness. These results suggest that reinforcement learning, when combined with transfer-aware observation design and deployment mechanisms, provides a practical approach for dynamic ball-balancing control on quadruped robots.

View source

Similar papers

Jul 2026

WARL: Wrench-Augmented Reinforcement Learning for Task-Agnostic Learning in Legged Robots

This study proposes a new method, Wrench-Augmented Reinforcement Learning (WARL), which introduces a wrenche (force and torque) into the action space, and shows that introducing a wrench can encourage behaviors that do not sufficiently exploit the robot's physical embodiment.

Keita Yoneda, Kento Kawaharazuka, Kei Okada · 0 citations
Preprint Aug 2026

Learning Highly Dynamic Skills Transition for Quadruped Jumping Through Constrained Space

This work proposes a hierarchical reinforcement learning pipeline that empowers the robots to perform aggressive locomotion through constrained obstacles--a narrow gate, extending the lifelike agility of legged robots to match that of their biological counterparts.

Zeren Luo, Jiahui Zhang, Yimin Han et al. · 1 citation
Preprint Aug 2026

Spatiotemporal Agility: Time-Constrained Reinforcement Learning for Vision-Guided Dynamic Quadrupedal Interception

An integrated framework that combines a vision module for landing point and time prediction with a direct position and time conditioned RL locomotion policy, instead of intermediate velocity commands is proposed, which mitigates perception latency during dynamic interception.

Yi-Dong Zhu, Zibo Dai, Tong-Ning Zhang et al. · 0 citations
Sep 2026

PRVR: Learning Posture Recovery for a Legged Spherical Robot Using Vector Reward

Legged spherical robots feature both walking and rolling modes and they are useful in explorations, but often end up in various body-inverted postures after rolling and accidental falls. Posture recovery in such cases where all legs are in the air and lose leg-ground contacts is an open and challenging problem. We propose a posture recovery learning method using vector reward called PRVR. A traditional scalar reward cannot evaluate each element of a vector action. Here, a vector reward method is proposed to evaluate each element of a vector action, such that the action of each joint closely conforms to the expected behavior. To reduce learning complexity, a rotation symmetry-based training and deployment method is proposed. Body-inverted postures can be divided into several symmetric classes, and then through training on one class of body-inverted postures, the robot can recover from various body-inverted postures. Simulations and experiments on a six-legged spherical robot are used to verify the effectiveness.

Jinkai Wang, Xin Xu, Chenkun Qi et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.