LDI-3DGS: Structurally Robust 3-D Gaussian Splatting Driven by UAV-Borne LiDAR Depth and Intensity Priors
Abstract
3-D Gaussian splatting (3DGS) has demonstrated outstanding performance in novel view synthesis and 3-D scene reconstruction. While 3DGS can technically be optimized from random initialization, achieving high-quality and scale-accurate reconstruction in large-scale scenes still relies heavily on sparse structure-from-motion (SfM) point clouds for reliable initialization. In these expansive urban environments featuring weak-texture regions, purely visual SfM frequently fails to provide robust and scale-accurate geometric priors. Consequently, Gaussian primitives lack sufficient high-frequency visual constraints, causing them to undergo unconstrained splitting and erroneous cloning. This often leads to severe structural collapse, floaters, and rendering blur. To overcome these limitations, we propose LDI-3DGS, a structurally robust scene reconstruction pipeline driven by dual physical priors. First, a light detection and ranging (LiDAR)-constrained joint bundle adjustment is formulated to rigorously register images into a drift-free absolute physical coordinate system, effectively eliminating scale drift. Second, dense spatial depth maps serve as geometric anchors to pull the Gaussian field onto actual physical surfaces. Furthermore, to address the inherent rendering blur in weak-texture regions, we innovatively leverage LiDAR reflection intensity to resolve the shape-radiance ambiguity. By formulating a joint color-intensity radiance field and enforcing an intensity-guided pruning constraint, our framework introduces a robust physical material prior to suppress erroneous Gaussian cloning. Extensive experiments on real-world urban datasets demonstrate that LDI-3DGS better preserves geometric and textural details, including road markings and highly reflective structures. Moreover, our framework improves absolute 3-D geometric accuracy and $F1$ -scores while reducing the total number of Gaussian primitives by approximately 23% relative to the baseline, ultimately delivering a structurally robust, compact, and physically accurate 3-D scene representation.