USR-Drive is proposed, a unified conditional generative framework that, given only posed multi-view driving videos, jointly recovers dense dynamic geometry and instance-level object layouts within a shared scene representation and delivers state-of-the-art results for both dynamic reconstruction and 3D detection on the nuScenes and VKitti datasets.
Abstract
Spatial representation learning for autonomous driving aims to map raw visual signals into structured 3D scene representations, where object-centric bounding boxes and rendering-oriented 3D primitives (\eg, 3D Gaussians) serve as two distinct yet highly complementary levels for scene understanding. Existing methods typically treat dynamic reconstruction and instance-level perception as separate tasks, despite their shared goal of estimating the underlying 3D world state. As a result, dynamic reconstruction is under-constrained while 3D detection lacks geometric grounding. To address this gap, we propose USR-Drive, a unified conditional generative framework that, given only posed multi-view driving videos, jointly recovers dense dynamic geometry and instance-level object layouts within a shared scene representation. Specifically, USR-Drive represents dense Gaussian primitives and sparse 3D bounding boxes as two aligned latent token streams and jointly denoises them with a unified multi-modal diffusion Transformer. Unlike prior paradigms that use boxes as external conditions or predict them with detached modules, USR-Drive treats them as mutually constrained state variables with a Unified Positional Encoding (UPE) that aligns heterogeneous tokens within a shared metric spatiotemporal coordinate. Via such unified representation and generative framework, the two modalities reinforce each other: geometry supplies dense metric evidence for box prediction, while boxes provide instance-level structural priors that help preserve spatial consistency and reduce ambiguity in sequential 3D geometric representation. Our approach successfully delivers state-of-the-art results for both dynamic reconstruction and 3D detection on the nuScenes and VKitti datasets.
A Geometry-grounded Unified 3D Perception (GeoUP) framework that adapts the reconstruction-oriented latent of VGGT to calibrated, streaming multi-camera driving scenes and achieves SOTA performance across detection, occupancy, and depth estimation is presented.
Longfei Xu, Xiao-Hui Wang, Ze-Hao Huang et al.· 2 citations
This work proposes a foundation-feature Gaussian driving world model that unifies scene understanding, language-grounded reasoning, controllable 4D editing, and multi-modal generation within a single framework and introduces a foundation-feature Gaussian tokenizer that directly distills Qwen/SigLIP visual-language feat...
Tianchen Deng, Xue-Feng Chen, Shuang Wu et al.· 0 citations
SpatialCrafter is presented, a novel two-stage framework that addresses explorable image-to-scene generation issues by introducing a global 3D proxy for high-fidelity image-to-scene generation and appearance refinement and introduces Parallel Geometry Injection and Proxy-Aware Corruption training strategies.
Chuan Fang, Lingteng Qiu, Yixun Liang et al.· 1 citation
This work introduces SPAR3S, a sparse voxel-aligned 3D latent generative model for conditional scene completion without requiring ground-truth 3D data for supervision, and trains a masked autoregressive transformer that jointly models voxel occupancy and latent token values, enabling efficient and spatially consistent...
Thomas Lucas, Maxime Pietrantoni, Philippe Weinzaepfel et al.· 1 citation
Recent depth foundation models like Depth Anything 3 (DA3) achieve remarkable multi-view depth estimation but assume static 3D scenes, limiting their applicability to real-world dynamic environments. Existing training-free 4D methods like Easi3R and VGGT4D rely on correspondence-trained backbones whose attention encode...
Xin-Hao Xiang, Wei-Yang Li, Zhi-Jie Zheng et al.· 0 citations
PoseAdapter, a lightweight framework for high-fidelity 2.5D controllable image generation, and a Context-Aware Dual-Stream Representation, to resolve the generative trade-off between strict instance isolation and global coherence.
Yu-Feng Chi, Hui-Min Ma, Fan Gao et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.