This paper introduces a new Kinematic Normalizing Flow (KNF) model, trained on a large-scale kinematic dataset, that generates diverse yet feasible partial reference motions that effectively addresses whole-body control challenges by decomposing the complex loco-manipulation problem into partial reference motion generation and low-level imitation control.
Abstract
Loco-manipulation has recently shown promising capabilities; however, achieving high-precision control, managing the high-dimensional action space induced by many degrees of freedom (DoFs), and fully exploiting the inherent redundancy of whole-body systems remain challenging. In this paper, we propose a novel whole-body control framework that effectively addresses these challenges by decomposing the complex loco-manipulation problem into partial reference motion generation and low-level imitation control. We introduce a new Kinematic Normalizing Flow (KNF) model, trained on a large-scale kinematic dataset, that generates diverse yet feasible partial reference motions. A high-level controller is then trained to navigate the KNF's latent space to exploit redundant solutions, while a low-level controller ensures physically feasible and accurate motion execution. We validate our approach on the quadrupedal robot equipped with a six-DoF robotic arm. In simulation, experimental results show that our approach significantly outperforms state-of-the-art methods in terms of tracking accuracy and feasible workspace coverage. For hardware deployment, we evaluate the system over 24 episodes across 8 different mobile loco-manipulation tasks. The system achieves end-effector pose-tracking errors of 4.5 cm and 0.14 rad, while maintaining accurate locomotion tracking with linear and angular velocity errors of 0.1 m/s and 0.01 rad/s, respectively, outperforming competitive baselines. Our method represents a practical and powerful solution for accurate and generalized whole-body loco-manipulation in high-DoF robotic systems, with promising potential for diverse downstream robotic tasks.
3D imitation learning has demonstrated capability in diverse visuomotor tasks but often struggles with small objects or high-precision manipulation due to geometric sparsity. To address this, we propose the Multimodal Object-aware Policy (MOP), a novel framework that incorporates a geometry-guided fusion module to adaptively integrate 2D semantic features with 3D geometry for precise control. Additionally, we introduce a lightweight Object Position Prediction (OPP) module that serves as an auxiliary supervision signal. The training labels for this module are generated by using a cost-effective vision-based software tracker, effectively replacing expensive hardware-based motion capture systems. We evaluate our policy across 56 tasks on 3 simulation benchmarks. Experimental results demonstrate that MOP significantly outperforms baselines, achieving higher success rates with lower variance and efficient inference. Real-world experiments on 4 tasks further validate the robustness and transferability of our approach.
Keyu Zhang, Yicheng Yang, Lifeng Wang et al.· IEEE Robotics and Automation...· 0 citations
This work uses Sample-based Model Predictive Control entirely in simulation as an automated, rapidly tunable expert to generate massive offline datasets and validate the robustness of this sim-to-real framework by successfully deploying complex loco-manipulation skills across different morphologies.
Martin Schuck, Maks Sorokin, S. Manni et al.· 0 citations
A framework that distills privileged teacher policies into vision-based humanoid controllers via world-model-assisted distillation, and introduces Performance-Conditioned Guidance (PCG), a reward-driven adaptive distillation schedule that computes performance scores for both teacher and student to dynamically balance guidance and exploration.
HAF (Humanoid Adaptation Framework), a two-part framework consisting of HAF-VLA and HAF-Steer that transfers off-the-shelf generalist VLA foundation models to humanoid whole-body loco-manipulation, surpasses vanilla single-stage VLA baselines and improves whole-body coordination and task performance.
Langzhe Gu, Chengkai Hou, Meng Li et al.· 0 citations
GeniWorld is presented, an interactive world model for robots that generalizes robustly across unseen scenarios by explicitly decoupling embodiment kinematics from environmental dynamics, and generates diverse manipulation trajectories within the world model, improving downstream policy performance and robustness in complex environments.
This thesis introduces local shape descriptors that allow grasp poses to transfer across object categories by exploiting shared geometric structure and proposes a potential-function-based framework for reactive motion generation, where neural fields model smooth energy functions whose gradients generate well-behaved vector fields for control.
A. Tekden· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.