It is shown that a flexible and expressive transformer, PointZero, outperforms prior methods on the same data and demonstrates the utility of the pre-training objective by post-training PointZero for two downstream applications: action-conditioned 3D dynamics prediction and imitation learning.
Abstract
World models endow perceptual systems with the ability to predict how scenes evolve under interaction. They are most beneficial when trained on diverse volumes of data, to instill a rich prior into downstream applications. Existing methods typically require robot action labels to learn action-conditioned 3D dynamics, which excludes web video data from the training pool. We study 3D point track completion as a pre-training objective for learning transferable 3D dynamics without robot data. Given a single RGB-D observation and sparse partial 3D trajectories (tracks), we predict future 3D tracks of all observed points. We show this objective produces a rich 3D dynamics prior, without requiring robot action labels. We contribute a diverse dataset of 2.9 million synthetic frames spanning deformable, articulated, and rigid objects, and use it to train PointZero. We show that a flexible and expressive transformer, PointZero, outperforms prior methods on the same data. We demonstrate the utility of our pre-training objective by post-training PointZero for two downstream applications: (1) action-conditioned 3D dynamics prediction and (2) imitation learning. When fine-tuned to condition on end-effector pose, PointZero outperforms the baselines on the recent PGND 3D dynamics benchmark. When fine-tuned to predict robot actions and 3D tracks, PointZero outperforms or matches the baselines on 6/7 simulated and real-world robot manipulation tasks. We furthermore evaluate training PointZero from scratch to isolate the benefits of our proposed architecture from those of our proposed pre-training objective and dataset. We release the dataset, checkpoints, and full training recipe.
This work studies how to learn anticipation of how objects move during human-object interaction from ordinary monocular videos of human-object interaction, and shows that it can be efficiently re-purpose video priors into explicit geometric forecasts for embodied intelligence.
GaussianDream demonstrates that training-time current Gaussian reconstruction and future Gaussian prediction provide effective 3D supervision, but its dense VGGT/TGE-based prefix jointly carries state, dynamics, and action-conditioning information.
Yu-Qing Jiang, Zi-Jian Zhang, Wei-Tao Zhou et al.· 0 citations
Action-conditioned video world models predict future robot interactions from multiple cameras, yet their outputs remain disparate video collections rather than a shared metric scene queryable across viewpoints and time. While existing 4D reconstruction methods offer a path to spatialize these predictions, independently...
Jin Hyun Kim, Min Young Kim, Soohwan Song et al.· 0 citations
PolyMem, an exemplar-free approach that implicitly models rich high-order statistics of the feature distribution to enhance cross-domain robustness, is introduced that effectively alleviates the performance discrepancy while improving the model's performance across domains.
Jin-Ge Ma, Gautham Vinod, Bruce Coburn et al.· 0 citations
Robot manipulation models primarily reason from 2D observations while acting in the 3D physical world. To bridge this gap, recent work has augmented robot data with geometric priors such as depth, point clouds, and 3D trajectories, while renderable 3D Gaussian representations provide another promising form of 3D superv...
World action models (WAMs) combine robot action generation with future state prediction. Existing WAMs typically predict videos or learned visual latents, which represent interaction geometry only implicitly and may retain appearance information unrelated to control. We introduce SkeleWAM, a compact WAM that represents...
Ju-Yi Sheng, Hua Wang, Meng-Yuan Liu· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.