Results indicate that a direct, multi-attribute 3D consistency objective, when combined with high-quality correspondences, is effective for addressing the ill-posed sparse-view reconstruction problem.
Abstract
Reconstructing high-fidelity 3D scenes from sparse-views remains a central problem in generalizable neural rendering. Existing generalizable 3D Gaussian Splatting (3DGS) methods often exhibit geometric artifacts in sparse-view settings, since supervision based solely on 2D photometric losses cannot resolve depth and correspondence ambiguities. To address this issue, we propose MAC-Splat, a training framework built around direct 3D consistency supervision. MAC-Splat builds on the MASt3R geometric backbone and a frozen DINOv3 encoder to obtain semantically informed 2D correspondences, which serve as geometric anchors for 3D supervision. Using these anchors, we define the Multi-Attribute Consistency (MAC) loss. This objective jointly regularizes the 3D attributes of matched Gaussians, including their position, shape, and appearance, by enforcing agreement in a common world coordinate frame. The formulation is robust to outliers and respects the geometry of covariance matrices, which leads to stable training under sparse-view conditions. Experiments on ScanNet++ show that MAC-Splat outperforms strong baselines, with particularly large gains under different overlap regimes. In particular, it improves average PSNR over Splatt3R by more than 4.5 dB, reduces LPIPS, and maintains performance as the camera pose gap increases. These results indicate that a direct, multi-attribute 3D consistency objective, when combined with high-quality correspondences, is effective for addressing the ill-posed sparse-view reconstruction problem.
A semantic-guided 3D Gaussian splatting (3DGS) framework tailored to sparse-view industrial reconstruction was introduced, enabling robust reconstruction from limited viewpoints and offers a practical geometric foundation for automated inspection and remote equipment monitoring.
Boyang Li, Tian-Han Gao, Zuan Gu et al.· Visual Computing for Industr...· 0 citations
Few-view surface reconstruction recovers the visible surfaces of a scene from a few posed RGB images, providing the 3D models that robots need to explore and interact online. On mobile platforms, the reconstruction must be fast and geometrically accurate while keeping a small memory footprint to ensure safe and efficient operation. 3D Gaussian Splatting (3DGS) offers a high-fidelity scene representation, but building it from a few views is ill-posed, as many distinct surfaces reproduce the same images, making traditional photometric methods prone to"floater"artifacts. End-to-end methods resolve the ambiguity by regressing splats with large, usually Transformer-based, networks that require heavy compute and memory while generalizing poorly to new scenes. We propose G2SR, which exploits a well-posed core of the task: given cross-view 2D splat correspondences, 3D splats follow analytically from multi-view geometry. G2SR employs a lightweight neural frontend to detect and track 2D Gaussian splats on the image plane and an analytic backend to triangulate each into a metric-scale 3D splat. On ScanNet, Replica, and DTU, G2SR matches or exceeds the geometric accuracy of state-of-the-art end-to-end methods while running at 69-89 reconstructions per second within 203 MB of GPU memory (5-107x less) for 2- and 3-view inputs at 384 x 512 resolution, offering a practical path to online Gaussian-based surface reconstruction.
Dasong Gao, Vivienne Sze, S. Karaman· arXiv.org· 0 citations
High-fidelity novel view synthesis from sparse inputs remains a significant challenge in 3D reconstruction. Traditional methods relying on Multi-layer Perceptrons (MLPs) often struggle with overfitting and inaccurate geometric representations, especially with limited training views. To address these issues, we introduce OC-CDGS, a novel framework leveraging Octree-structured 3D Gaussian Splatting (3DGS) with Combined Depth Regularization (CDR). OC-CDGS organizes sparse point clouds into hierarchical octree-structured anchors, enabling efficient decoding of neural Gaussian attributes through compact MLPs. The CDR approach integrates global structure, local details, and gradient consistency to enhance geometric supervision. Additionally, we incorporate a Neural Gaussian Random Dropout Strategy (NGDS) to improve representation robustness and a Scene-Extent- Based Anchor Densification Strategy (SADS) to enhance anchor coverage in peripheral regions. Extensive experiments on MipNeRF360 and LLFF datasets demonstrate that OC-CDGS achieves state-of-the-art performance with real-time rendering capabilities, preserving high-quality geometric details.
A 3D-GS-based framework that attaches a simplified Disney Bidirectional Reflectance Distribution Function (BRDF) and differentiable physically based rendering (PBR) shader to each surface Gaussian for interpretable material attributes delivers higher geometric accuracy, improved material realism, and real-time rendering performance suitable for applications such as Simultaneous Localization and Mapping (SLAM) and relighting.
Wei-Chen Xu, W. Wan, M. Peng· International Conference on...· 0 citations
This work proposes a novel 3D-aware video restoration framework designed to enhance the quality of sparse 3DGS reconstruction and introduces a camera-conditioned geometric prior that guides the network toward geometrically grounded restoration that remains coherent across viewpoints.
Xinhui Liu, Can Wang, Wei Jiang et al.· 0 citations
3D Gaussian Splatting (3DGS) delivers real-time and high-fidelity rendering but remains challenged by unconstrained in-the-wild scenes, where drastic appearance variations and transient objects violate multi-view consistency. Existing methods are fundamentally limited by independent and discrete embeddings that struggle to capture continuous environmental changes or model spatially-varying local illumination. To address these limitations, we propose \textbf{WilLaGS}, a unified framework for robust 3D scene reconstruction and generative appearance synthesis under unconstrained settings. Specifically, we introduce a generative appearance model where a $\beta$-VAE learns a structured and continuous manifold of global appearance. Conditioned on the latent code, we construct a 3D neural appearance field that generates dynamic Tri-Plane features to encode spatially-varying local illumination effects. Furthermore, to suppress transient artifacts, we present a self-supervised perceptual masking mechanism that leverages a Teacher-Student (EMA) architecture to derive a stable scene consensus, robustly identifying inconsistent regions via perceptual discrepancies. Extensive experiments on multiple datasets demonstrate that \textbf{WilLaGS} achieves state-of-the-art performance in reconstruction quality and novel view appearance synthesis, while maintaining real-time rendering efficiency.
Yu Bai, Qian-Qiu Tan, Li-Long Chen et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.