Skip to content
Preprint

GenTrack: Physical Alignment for Robot-Native Motion Generation and Zero-Shot Humanoid Tracking

Aug 2026 · 0 citations · 55 references
Computer Science

TL;DR

GenTrack is introduced, an online generator--tracker framework that alternates execution-grounded, group-relative generator alignment with tracker training on newly generated references; anchoring and rehearsal constrain drift and demonstrates that joint online post-training effectively narrows the executability gap between retargeted references and robot-native motion.

Abstract

General-purpose humanoid trackers can execute diverse references, but their zero-shot coverage depends on large embodied corpora that are costly to extend. Text-to-motion generators offer scalable supervision, yet models trained on human motion or retargeted data inherit a gap between kinematic plausibility and robot executability. Existing one-way pipelines fix either the generated corpus or the reward tracker. We introduce GenTrack, an online generator--tracker framework that alternates execution-grounded, group-relative generator alignment with tracker training on newly generated references; anchoring and rehearsal constrain drift. On Unitree G1, we evaluate GenTrack with ProtoMotions and SONIC backbones across three zero-shot tracking splits including public AMASS and LAFAN benchmarks, and a private out-of-distribution test set of 1,024 prompt-motion pairs in the wild. The online co-training strategy consistently produces generators that output more robot-executable motions with strong semantic alignment, and trackers with markedly broader zero-shot coverage and improved tracking accuracy, especially on out-of-distribution references. These results demonstrate that joint online post-training effectively narrows the executability gap between retargeted references and robot-native motion, advancing zero-shot humanoid control without additional data collection and beyond the limitations of a static reference pool.

View source

Similar papers

Preprint Aug 2026

HumanTracker: Towards Comprehensive and Human-Aligned Motion Tracking Benchmark

This work introduces HumanTracker, a preference-aligned metric trained on 12K motion pairs containing 24K motions that better predicts human preferences and reveals contact and stability failures that kinematic metrics often miss.

Dai-En Liu, Zekun Qi, Jiayu Zeng et al. · 0 citations
Preprint Aug 2026

StableMimic: Smooth Human-Like Recovery for Humanoid Motion Tracking - Learning Beyond the Tracking Distribution for Structured Post-Fall Behavior

StableMimic is presented, a unified tracker trained beyond the nominal tracking distribution that achieves the lowest errors on all four tracking metrics among five methods and attains the lowest values on six of seven post-fall motion and load measures, supporting improved interaction safety under this protocol.

Weihao Wu, Mingzhe Huang, Ruofei Liu et al. · 0 citations
Jul 2026

What Matters in Humanoid General Motion Tracking? An Empirical Study

An empirical study of common modeling and training factors used in recent humanoid motion-imitation pipelines, developing YAHMP, an open-source modular framework for training, evaluating, and deploying whole-body motion tracking policies on the Unitree G1.

Fabio Amadio, Enrico Mingo Hoffman · 1 citation

Lightweight Adaptation of Pretrained Robot Manipulation Systems: Two Approaches

Two systematic attempts to improve large pretrained models with minimal or zero modification to their weights via reinforcement learning on a frozen OpenVLA-7B using binary task-success rewards on LIBERO-Goal reveal a common ceiling.

Adam Lalani, Chen Sun, Hui Wang · 0 citations
Review Aug 2026

HumanoidVLN: A Physics-Grounded Simulator and Benchmark for Vision-Language Navigation Across Diverse Humanoid Embodiments

Vision-Language Navigation (VLN) for humanoid robots poses challenges existing benchmarks fail to address: bipedal locomotion imposes physical constraints absent from wheeled agents, humanoid morphologies vary across platforms, and egocentric observations are distorted by locomotion-induced camera dynamics. We present HumanoidVLN, a physics-grounded simulator and benchmark for VLN across diverse humanoid embodiments. Built on NVIDIA Isaac Sim, our platform supports an extensible set of humanoid configurations, demonstrated on four robots (Unitree G1, Unitree H1, Internal-A, Internal-B) spanning 10-12 lower-body DoF and heights from 1.17m to 1.80m, via a hierarchical control stack combining a reinforcement learning locomotion policy with interchangeable PD or MPC path trackers. New robots and VLN models integrate with minimal effort; we demonstrate compatibility with NaVILA, DualVLN, StreamVLN, and JanusVLN. Environments are drawn from artist-designed scenes and 3D Gaussian Splatting reconstructions, filtered for navigable areas exceeding 100 square meters. Instructions are generated by a dual generator-reviewer plus paraphraser multi-agent pipeline with human-in-the-loop verification, yielding 933 collision-aware reference episodes, each paired with one fine-grained instruction and three coarse-grained stylistic variants (formal, natural, casual). Across four models and four embodiments, JanusVLN achieves the highest mean success rate of 43.55% and nDTW of 48.38. In a 20-episode sim-to-real pilot with DualVLN and the Unitree G1, navigation errors correlate strongly (r=0.935), with a mean absolute difference of 0.68m and mean trajectory similarity of 0.782 (+/-0.188) nDTW. These results highlight the interaction between VLN models, controllers, and humanoid embodiments under physical execution. Code, benchmark, and data will be released upon acceptance at https://humanoid-vln.github.io/.

Q. Phạm, Anh Dao, T. Nguyen et al. · 0 citations
Preprint Aug 2026

RoboReact: Agentic Skill Distillation from Generated Egocentric Videos for Generalizable Whole-Body Manipulation

This work presents RoboReact, a framework that automatically synthesizes whole-body humanoid manipulation skills from a single egocentric RGB-D observation, and highlights the potential of combining generative models, vision-language reasoning, and closed-loop control for scalable humanoid skill acquisition.

Shuliang He, Shuai Wang, Bo Yue et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.