M3GD is presented, a Camera--LiDAR multimodal representation for generative NVS that composes independently pretrained 2D image and 3D point-cloud foundation models without separately pretraining a cross-modal translator.
Abstract
Robotic novel view synthesis (NVS) must recover both visual appearance and metric 3D structure, yet most generative NVS methods rely only on images, overlooking LiDAR, a complementary sensor common on robotic platforms. We present M3GD, a Camera--LiDAR multimodal representation for generative NVS that composes independently pretrained 2D image and 3D point-cloud foundation models without separately pretraining a cross-modal translator. We show that, after camera projection, frozen LiDAR and image features exhibit substantial shared spatial structure, providing a natural cross-modal representation. M3GD conditions generation on LiDAR through this structure: it combines explicit geometry statistics with learned point-cloud descriptors into view-aligned packets on the image-latent grid, injected through a lightweight residual adapter into a multi-view flow-matching generator whose latent space, decoders, and training objective remain intact. On the GrandTour dataset, M3GD improves target-view RGB and depth synthesis over an image-only version of the same backbone. Ablations show that the gains come from pixel-aligned LiDAR content and that target-view LiDAR acts as a geometric query linking the requested view to source observations. Deployment on a ground robot demonstrates practical real-world operation, with a configurable quality--cost trade-off controlled by the number of Euler integration steps.
We present DXPR, a depth-based cross-modal place recognition (CMPR) framework that uses vision foundation models (VFMs) to match monocular camera queries against a LiDAR map without modality-specific encoders. This enables robots and autonomous vehicles to robustly localize using only cameras within pre-built LiDAR map...
Yu-Hang Han, Youngseok Jang, Seungwon Roh et al.· 0 citations
Synthesizing novel camera views is important for autonomous ground vehicles, with applications in surround-view monitoring, occlusion recovery, and training data augmentation. We present View Translation, a geometry-guided latent diffusion framework that generates a target camera view from a source image, relative came...
O. Mayekar, M. Aiyetigbo, Ameya Salvi et al.· SAE technical paper series· 0 citations
Cyclops is proposed, a framework that translates sparse Non-Repetitive Scanning LiDAR intensity into RGB video, enabling camera-free inference for all-day perception tasks and mitigating inter-frame flickering.
Wei Gao, Jian Shu, Ming-Le Zhao et al.· 0 citations
A Geometry-grounded Unified 3D Perception (GeoUP) framework that adapts the reconstruction-oriented latent of VGGT to calibrated, streaming multi-camera driving scenes and achieves SOTA performance across detection, occupancy, and depth estimation is presented.
Longfei Xu, Xiao-Hui Wang, Ze-Hao Huang et al.· 2 citations
Non-contact 3D measurements acquired by LiDAR, structured-light and depth-camera systems often contain large missing regions under occlusion, limited viewpoints, motion blur and sensor noise. These defects reduce the geometric fidelity of reconstructed shapes and directly affect downstream dimensional analysis, pose es...
Yu-Hao Yang, Gun Li, Jia-Cheng Luo et al.· Measurement science and tech...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.