Skip to content

Author

Shuhan Shen

2 papers indexed here

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

Jul 2026

M4World: A Multi-view Multimodal Driving World Model for Interactive Object Manipulation and Minute-long Streaming

Driving-world generation has emerged as a core capability for scalable autonomous-driving simulation, yet existing methods remain limited in object-level controllability and long-horizon stability. We present M$^\text{4}$World, a Multi-view and Multimodal generative driving world model that synthesizes future surround-view video streams and synchronized LiDAR scans while supporting interactive object Manipulation and stable Minute-long streaming. Fine-grained object manipulation is realized through a flexible conditioning interface that supports explicit control over both the spatial layout and visual appearance of individual objects. Stable minute-long streaming, on the other hand, is achieved through a multi-stage training framework that enables online causal generation in only four denoising steps while maintaining coherent world dynamics throughout extended rollouts. Building on these components, we introduce an efficient few-clip post-training as well as a suite of visual reference-conditioned generation models, preserving general generation ability while allowing rare-case customization for long-tail controllability. To assess controllability beyond realism, we further introduce an automated VLM-based judging pipeline that evaluates scene-level condition adherence, view-wise object controllability, and cross-view object consistency. Comprehensive experiments show that M$^\text{4}$World consistently delivers high generation quality, precise controllability, and stable minute-long streaming. Together with downstream long-tail augmentation and scene editing, these results demonstrate the potential of M$^\text{4}$World for controllable, scalable driving simulation.

Ke Cheng, Hanqiao Ye, Lei Shi et al. · 0 citations
2026

VLGS-SLAM: Visual–Lidar Three-Dimensional Gaussian Splatting Simultaneous Localization and Mapping

Recent studies highlight the effectiveness of 3D Gaussian splatting (3DGS) in visual simultaneous localization and mapping (SLAM) systems, which work well indoors but struggle in large-scale outdoor environments. Typically, lidar data are used to address this issue; however, current multi-modal SLAM systems use 3DGS mainly for mapping, leaving its potential to enhance tracking unexplored. In this paper, we present VLGS-SLAM, a novel visual–lidar SLAM pipeline that leverages lidar data and 3DGS for both pose estimation and mapping. Our approach integrates lidar points as 3D Gaussian primitives, ensuring precise scene geometry and reducing pose estimation errors caused by floating Gaussians. To enhance tracking performance, we apply regularization to Gaussian scaling, which constrains the shape of each Gaussian ellipsoid. For loop closure, we combine image similarity with lidar cloud distance to effectively detect and close loops. Our experiments demonstrate that VLGS-SLAM achieves state-of-the-art accuracy in the 3DGS-based SLAM field, outperforming many traditional SLAM algorithms

Diantao Tu, Wen-Juan Ma, Shuhan Shen · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.