Skip to content
Preprint

PointZero: 3D Point Track Completion for Learning Transferable 3D Dynamics

Sep 2026 · 0 citations · 45 references
Computer Science

TL;DR

It is shown that a flexible and expressive transformer, PointZero, outperforms prior methods on the same data and demonstrates the utility of the pre-training objective by post-training PointZero for two downstream applications: action-conditioned 3D dynamics prediction and imitation learning.

Abstract

World models endow perceptual systems with the ability to predict how scenes evolve under interaction. They are most beneficial when trained on diverse volumes of data, to instill a rich prior into downstream applications. Existing methods typically require robot action labels to learn action-conditioned 3D dynamics, which excludes web video data from the training pool. We study 3D point track completion as a pre-training objective for learning transferable 3D dynamics without robot data. Given a single RGB-D observation and sparse partial 3D trajectories (tracks), we predict future 3D tracks of all observed points. We show this objective produces a rich 3D dynamics prior, without requiring robot action labels. We contribute a diverse dataset of 2.9 million synthetic frames spanning deformable, articulated, and rigid objects, and use it to train PointZero. We show that a flexible and expressive transformer, PointZero, outperforms prior methods on the same data. We demonstrate the utility of our pre-training objective by post-training PointZero for two downstream applications: (1) action-conditioned 3D dynamics prediction and (2) imitation learning. When fine-tuned to condition on end-effector pose, PointZero outperforms the baselines on the recent PGND 3D dynamics benchmark. When fine-tuned to predict robot actions and 3D tracks, PointZero outperforms or matches the baselines on 6/7 simulated and real-world robot manipulation tasks. We furthermore evaluate training PointZero from scratch to isolate the benefits of our proposed architecture from those of our proposed pre-training objective and dataset. We release the dataset, checkpoints, and full training recipe.

View source

Similar papers

MotionForesight Re-purposing Video Models for Future 3D Scene-Flow Prediction

This work studies how to learn anticipation of how objects move during human-object interaction from ordinary monocular videos of human-object interaction, and shows that it can be efficiently re-purpose video priors into explicit geometric forecasts for embodied intelligence.

Unknown authors · 0 citations
Preprint Aug 2026

GaussianDream++: Efficient 3D Gaussian World Modeling for Robotic Manipulation

GaussianDream demonstrates that training-time current Gaussian reconstruction and future Gaussian prediction provide effective 3D supervision, but its dense VGGT/TGE-based prefix jointly carries state, dynamics, and action-conditioning information.

Yu-Qing Jiang, Zi-Jian Zhang, Wei-Tao Zhou et al. · 0 citations
Preprint Sep 2026

RoGSW4RLD: Feed-Forward 4D Gaussian Lifting for Robot World Model Rollouts

Action-conditioned video world models predict future robot interactions from multiple cameras, yet their outputs remain disparate video collections rather than a shared metric scene queryable across viewpoints and time. While existing 4D reconstruction methods offer a path to spatialize these predictions, independently...

Jin Hyun Kim, Min Young Kim, Soohwan Song et al. · 0 citations
#artificial intelligence Preprint Sep 2026

Mitigating Performance Discrepancy in Cross-Domain 3D Class-Incremental Learning

PolyMem, an exemplar-free approach that implicitly models rich high-order statistics of the feature distribution to enhance cross-domain robustness, is introduced that effectively alleviates the performance discrepancy while improving the model's performance across domains.

Jin-Ge Ma, Gautham Vinod, Bruce Coburn et al. · 0 citations
Preprint Oct 2026

3DROID: A Renderable 3D Gaussian Dataset with Measured Per-Scene Reliability

Robot manipulation models primarily reason from 2D observations while acting in the 3D physical world. To bridge this gap, recent work has augmented robot data with geometric priors such as depth, point clouds, and 3D trajectories, while renderable 3D Gaussian representations provide another promising form of 3D superv...

W. Cho, Junhoo Lee, N. Kwak · 0 citations
Preprint Oct 2026

SkeleWAM: Skeleton World-Action Modeling for Efficient Robotic Manipulation

World action models (WAMs) combine robot action generation with future state prediction. Existing WAMs typically predict videos or learned visual latents, which represent interaction geometry only implicitly and may retain appearance information unrelated to control. We introduce SkeleWAM, a compact WAM that represents...

Ju-Yi Sheng, Hua Wang, Meng-Yuan Liu · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.