World-action models (WAMs) jointly model future observations and actions, typically predicting the future as RGB images. Other visual modalities such as depth, pretrained visual features, and point tracks can more efficiently capture geometric, semantic, and motion features. However, how best to combine these modalitie...
Adam Hung, B. Duisterhof, D. Ramanan et al.· 1 citation
It is shown that a flexible and expressive transformer, PointZero, outperforms prior methods on the same data and demonstrates the utility of the pre-training objective by post-training PointZero for two downstream applications: action-conditioned 3D dynamics prediction and imitation learning.
B. Duisterhof, Kai-Feng Zhang, Adam Hung et al.· 0 citations
This work introduces SPAR3S, a sparse voxel-aligned 3D latent generative model for conditional scene completion without requiring ground-truth 3D data for supervision, and trains a masked autoregressive transformer that jointly models voxel occupancy and latent token values, enabling efficient and spatially consistent...
Thomas Lucas, Maxime Pietrantoni, Philippe Weinzaepfel et al.· 1 citation
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.