Skip to content

Enhance 3D Gaussian splatting for dynamic scenes: integrating semantic and geometric consistency

Jul 2026 · Engineering Research Express · Vol 8, pp. 145202 · 0 citations · 32 references
Physics

TL;DR

The proposed approach achieves competitive or superior performance compared with 3DGS, SpotLessSplats, T-3DGS, and RobustSplat on standard image-quality metrics, including peak signal-to-noise ratio, structural similarity index measure, and learned perceptual image patch similarity.

Abstract

3D Gaussian splatting (3DGS) provides an efficient and expressive scene representation by jointly modeling spatial geometry and appearance, which has led to significant advances in high-fidelity scene reconstruction and novel view synthesis. However, in real-world environments, occlusions from pedestrians or equipment often introduce blurriness, artifacts, and geometric distortions. To address these challenges, this paper proposes a robust 3DGS modeling method enhanced by semantic and geometric consistency. First, the self-supervised foundation model DINOv2 is utilized to extract high-dimensional semantic features, leveraging its superior generalization capabilities to assist in identifying potential dynamic regions. Second, monocular depth estimation and a depth residual mechanism are introduced to construct geometric consistency constraints, enabling the precise localization of areas that violate static assumptions. Finally, a progressive guided probability masking mechanism is designed; it employs an adaptive sigmoid function to achieve a ‘coarse-to-fine’ soft-constraint optimization, effectively mitigating the training instability inherent in traditional binary hard masks. Experimental results on the neural radiance fields (NeRF)-on-the-go, RobustNeRF, and self-collected datasets demonstrate that the proposed method effectively suppresses dynamic artifacts and improves reconstruction quality. The proposed approach achieves competitive or superior performance compared with 3DGS, SpotLessSplats, T-3DGS, and RobustSplat on standard image-quality metrics, including peak signal-to-noise ratio, structural similarity index measure, and learned perceptual image patch similarity.

View source

Similar papers

Jul 2026

Geometry-Semantics Co-Regularization for Gaussian Splatting in Indoor Reconstruction.

A geometry-semantics co-regularization framework that jointly optimizes geometry and semantics within 3DGS and develops a multi-view semantic consistency supervision to regularize the semantic distributions of Gaussian primitives, ensuring cross-view consistency for Gaussians corresponding to the same semantic category or instance.

Haihong Xiao, Jianan Zou, Yanan Zhang et al. · 0 citations
Jul 2026

UH-GS: Uncertainty-aware hierarchical Gaussian splatting for outdoor scene reconstruction

Abstract. In outdoor scene reconstruction, dynamic occlusions and multiscale structures often undermine multiview consistency and hinder effective gradient accumulation of high-frequency Gaussian primitives, leading to artifacts and the loss of fine details in Gaussian splatting–based radiance field methods. To address these challenges, we propose an uncertainty-aware hierarchical Gaussian splatting framework for outdoor 3D reconstruction. Specifically, our method constructs a hierarchical octree-based spatial representation from the results of aerial triangulation. It introduces level of detail constraints to enable structured management and progressive optimization of Gaussian primitives across different scales. This design effectively alleviates the imbalance in training and the redundant growth of Gaussian primitives commonly observed in multiscale outdoor scenes. In addition, we incorporate an uncertainty prediction mechanism that evaluates the consistency between rendered results and ground-truth images in the feature space, allowing the model to automatically identify dynamically occluded regions and suppress their gradient contributions during optimization. As a result, the adverse impact of dynamic artifacts on static scene modeling is substantially reduced. Experimental results demonstrate that, without incurring significant additional training overhead, our method consistently improves structural consistency and fine-detail reconstruction quality in outdoor scenes, while simultaneously reducing model complexity and maintaining real-time rendering performance. Furthermore, the proposed approach can be seamlessly integrated into multiple mainstream Gaussian splatting frameworks, exhibiting strong robustness and promising potential for practical deployment.

Junxing Yang, Haoran Gao, Chun-Yu Huang et al. · 0 citations
Open access Aug 2026

Semantic-guided 3D Gaussian splatting for sparse-view reconstruction in industrial digital twins

A semantic-guided 3D Gaussian splatting (3DGS) framework tailored to sparse-view industrial reconstruction was introduced, enabling robust reconstruction from limited viewpoints and offers a practical geometric foundation for automated inspection and remote equipment monitoring.

Boyang Li, Tian-Han Gao, Zuan Gu et al. · 0 citations
Preprint Aug 2026

WilLaGS: Latent-Conditional 3D Appearance Fields for Robust Gaussian Splatting In-the-Wild

3D Gaussian Splatting (3DGS) delivers real-time and high-fidelity rendering but remains challenged by unconstrained in-the-wild scenes, where drastic appearance variations and transient objects violate multi-view consistency. Existing methods are fundamentally limited by independent and discrete embeddings that struggle to capture continuous environmental changes or model spatially-varying local illumination. To address these limitations, we propose \textbf{WilLaGS}, a unified framework for robust 3D scene reconstruction and generative appearance synthesis under unconstrained settings. Specifically, we introduce a generative appearance model where a $\beta$-VAE learns a structured and continuous manifold of global appearance. Conditioned on the latent code, we construct a 3D neural appearance field that generates dynamic Tri-Plane features to encode spatially-varying local illumination effects. Furthermore, to suppress transient artifacts, we present a self-supervised perceptual masking mechanism that leverages a Teacher-Student (EMA) architecture to derive a stable scene consensus, robustly identifying inconsistent regions via perceptual discrepancies. Extensive experiments on multiple datasets demonstrate that \textbf{WilLaGS} achieves state-of-the-art performance in reconstruction quality and novel view appearance synthesis, while maintaining real-time rendering efficiency.

Yu Bai, Qian-Qiu Tan, Li-Long Chen et al. · 0 citations
Jul 2026

Consistent 4D Appearance Editing with Gaussian Splatting.

Editing dynamic scenes with 4D Gaussian Splatting (4DGS) is often hampered by spatiotemporal inconsistencies, or "Gaussian drifting", which degrades edit quality and temporal coherence. We identify that these artifacts stem from two distinct sources: foundational inaccuracies in the initial scene reconstruction, and the disruption of learned trajectories during the editing process itself. To address this, we propose a comprehensive framework that systematically tackles both sources of inconsistency. To solve reconstruction-induced errors, we introduce a novel prior-guided, multi-stage reconstruction pipeline that fuses geometric and motion priors to build a physically plausible and temporally stable foundation. To solve editing-induced errors, we further apply a universal trajectory-preserving technique, which safeguards high-quality motion by decoupling the appearance optimization from the learned deformation. Experiments demonstrate that by systematically addressing both the reconstruction and editing phases, our method achieves state-of-the-art, temporally consistent editing on a wide range of dynamic scenes where previous monolithic approaches fail.

Xiaosheng He, Feng-Lin Liu, Lin Gao et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.