Skip to content
Preprint

Spatiotemporal Agility: Time-Constrained Reinforcement Learning for Vision-Guided Dynamic Quadrupedal Interception

Aug 2026 · 0 citations · 26 references
Computer Science

TL;DR

An integrated framework that combines a vision module for landing point and time prediction with a direct position and time conditioned RL locomotion policy, instead of intermediate velocity commands is proposed, which mitigates perception latency during dynamic interception.

Abstract

Legged robots require robust agility to perceive and interact with complex and dynamic environments within a constrained time. However, most existing quadruped locomotion works rely on velocity-tracking policy, which struggle to reach precise targets within strict temporal constraints. Moreover, integrating real-time perception with agile locomotion for highly dynamic targets remains challenging due to sensor latency and processing delays. To concretely study and benchmark such agility in dynamic settings, we introduce a challenging ball-catching task for legged robots. This paper proposes an integrated framework that combines a vision module for landing point and time prediction with a direct position and time conditioned RL locomotion policy, instead of intermediate velocity commands. Beyond the method design, this work presents a system-level contribution that completes real-time robotic interception system that integrates multi-camera perception, online trajectory prediction, low-latency target communication, and sim-to-real locomotion control into a closed-loop deployment pipeline. By explicitly predicting the future spatial-temporal target, our approach mitigates perception latency during dynamic interception. We conducted extensive ball-catching experiments for the legged robot. Through comparative experiments against a velocity-tracking baseline, our direct target-conditioned approach achieves a higher success rate in catching balls with predicted landing spots within 2 meters and flight times between 0.8 and 1.2 seconds. This shows that the robot has successfully completed the dynamic ball-catching task under our tested setup. Furthermore, our policy exhibits a smaller performance gap after deployment, suggesting improved sim-to-real behavior in these trials.

View source

Similar papers

Preprint Aug 2026

Learning Highly Dynamic Skills Transition for Quadruped Jumping Through Constrained Space

This work proposes a hierarchical reinforcement learning pipeline that empowers the robots to perform aggressive locomotion through constrained obstacles--a narrow gate, extending the lifelike agility of legged robots to match that of their biological counterparts.

Zeren Luo, Jiahui Zhang, Yimin Han et al. · 1 citation
Conference Aug 2026

Vision-Based Predictive Control for Dual-Arm Nonprehensile Transportation

Dual-arm robots often encounter difficulties when handling easily deformable or structurally complex objects using traditional grasping-based manipulation. In addition, grasping and releasing operations introduce significant time overhead. To address these limitations, this paper proposes a vision-based predictive control framework for dual-arm nonprehensile transportation. The proposed method employs a hybrid end effector design that integrates an elastic tether with a tray, enabling flexible and stable transportation without direct grasping. A predictive control strategy is adopted to optimize dual-arm motion trajectories on the move under kinematic and safety constraints. To further enhance coordination accuracy, a direct visual servoing scheme is incorporated to dynamically regulate the arm velocities, minimizing relative motion between the end effectors and the object. This effectively suppresses oscillations induced by the elastic tether. Both simulation and experimental results demonstrate that the proposed approach ensures convergence to desired states and achieves continuous, stable, and safe object transportation, even in the presence of disturbances.

Chang Liu, Yuan Yang, Panfeng Huang et al. · 0 citations
Preprint Aug 2026

Learning Fault-Tolerant Locomotion with Adaptive Gait Timing

Hardware failures require legged robots to rapidly reorganize coordination and gait timing to maintain stability and mobility. This is particularly challenging for larger quadrupeds, where increased mass and tighter actuation limits reduce the feasibility of aggressive, high-frequency compensation strategies often observed on smaller platforms. In this work, we propose a deep reinforcement learning approach for fault-tolerant locomotion under actuator power loss. The method employs an asymmetric actor-critic architecture in which the critic has access to privileged information during training, while the actor learns to reconstruct a corresponding latent representation from proprioceptive observations. We introduce a latent-alignment loss that encourages consistency between actor and critic representations. Additionally, we augment the action space with a learnable gait frequency parameter, enabling adaptive gait timing in response to terrain variations and actuator degradation without predefined faulty-leg strategies. The approach is validated in high-fidelity simulation on uneven terrain and real-world experiments on flat ground using a 68 kg quadruped robot.

Giovanbattista Gravina, Luca Rossini, Carlo Rizzardo et al. · 0 citations
Preprint Sep 2026

Learning-Based Dynamic Obstacle Avoidance for a UAV Using Only Three Range Sensors

We present a learning-based approach to kinodynamic online motion planning for an Unmanned Aerial Vehicle (UAV) operating at a fixed altitude in unknown dynamic environments, where real-time avoidance of both static and dynamic obstacles must be achieved under conditions of extreme partial observability. The UAV is controlled with a single degree of freedom (yaw only), resulting in constrained, nonholonomic motion similar to fixed-wing platforms. The proposed framework integrates a behavior grid map representation with Deep Reinforcement Learning (DRL), using Proximal Policy Optimization (PPO) for stable policy learning in continuous control. The key idea is the co-design of a state representation and control policy that enables reliable navigation using only three low-cost directional range sensors, without reliance on dense sensing modalities such as LiDAR or vision-based systems. The behavior grid map dynamically aggregates sparse measurements into a structured local representation that supports real-time decision-making for obstacle avoidance and target reaching. Extensive simulations across environments of varying sizes and obstacle densities demonstrate that the proposed standard and enhanced methods achieve higher success rates than PPO variants and Model Predictive Control (MPC) (94\% vs. 79--90\% in small-scale high-congestion scenarios, and 83\% vs. 62--71\% in large-scale high-congestion scenarios), while maintaining real-time performance. Real-world experiments across four scenarios further confirm practical feasibility, with consistent target-reaching behaviour and no collisions under the tested conditions.

Mohammad Reza Ranjbar Divkoti, A. Aguiar · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.