This paper proposes a framework for efficient incremental optimization of 3D Gaussian Splatting models and achieves a 24% relative improvement in SSIM with just 120 seconds of additional optimization on the Mip-NeRF360 dataset.
Abstract
Recent advances in 3D reconstruction using mobile cameras have expanded their potential applications beyond traditional domains such as surveying and virtual reality, extending to a wide range of industries. In particular, 3D Gaussian Splatting (3DGS), which can generate photorealistic novel view synthesis from video captured with off-the-shelf RGB cameras, shows promise for industrial use cases involving mobile camera systems that collect images in real time. However, pipelines based on offline Structure-from-Motion (SfM), e.g., COLMAP, are computationally expensive and thus limit practical deployment. While neural network–based acceleration methods have emerged, they are typically limited to processing a small number of input frames, constraining both reconstruction accuracy and spatial coverage. This paper proposes a framework for efficient incremental optimization of 3DGS models. Our framework enables fast and effective fine-tuning by adaptively adjusting the camera poses for additional image frames based on the relationship between L1 and SSIM rendering losses used for 3DGS optimization. Applied to rapidly initialized 3DGS models, our approach achieves a 24% relative improvement in SSIM with just 120 seconds of additional optimization on the Mip-NeRF360 dataset.
This work introduces FastEventDGS, a novel Deformable Gaussian Splatting-based framework that leverages a single event camera for high-fidelity 4D reconstruction in dynamic scenes and proposes a local patch event motion loss to constrain object motion, effectively mitigating over-fitting.
Zijia Dai, Nico Messikommer, Rong Zou et al.· 0 citations
A semantic-guided 3D Gaussian splatting (3DGS) framework tailored to sparse-view industrial reconstruction was introduced, enabling robust reconstruction from limited viewpoints and offers a practical geometric foundation for automated inspection and remote equipment monitoring.
Boyang Li, Tian-Han Gao, Zuan Gu et al.· Visual Computing for Industr...· 0 citations
Abstract. Monocular depth estimation (MDE) infers depth from a single image, offering significant advantages in computational efficiency and memory consumption compared to conventional Multi-View Stereo (MVS) methods. However, most MDE methods suffer from poor multi-view geometric consistency, which limits their application to photogrammetric 3D reconstruction. To address this issue, this paper employs sparse point clouds of Structure-from-Motion (SfM) as extra geometric constraints and proposes a framework that achieves photogrammetric 3D reconstruction using off-the-shelf learning-based MDE models without the need for additional fine-tuning. Specifically, when SfM priors are available during inference, globally geometrically consistent depth maps can be directly predicted. Otherwise, the estimated monocular depths are aligned to a consistent scale using SfM results via a post-correction step. The resulting depth maps are then fused using a truncated signed distance function (TSDF) to generate dense 3D reconstructions. Experiments on photogrammetric datasets demonstrate that the proposed framework effectively improves geometric consistency across depth maps and enables high-quality scene reconstruction. In addition, we systematically analyze the impact of key parameters in depth inference and fusion, including depth map resolution, voxel size, denoising steps, and ensemble size, on reconstruction performance, and further explore the potential of MDE for photogrammetric 3D reconstruction.
Chunyu Dou, Yifei Yu, Xin Wang et al.· The International Archives o...· 0 citations
Neural radiance fields (NeRF) and 3D Gaussian Splatting (3DGS) are popular techniques to reconstruct and render photorealistic images. However, the prerequisite of running Structure-from-Motion (SfM) to get camera poses limits their completeness. Although previous methods can reconstruct a few unposed images, they are not applicable when images are unordered or densely captured. In this work, we propose a method to train 3DGS from unposed images. Our method leverages a pre-trained 3D geometric foundation model as the neural scene representation. Since the accuracy of the predicted pointmaps does not suffice for accurate image registration and high-fidelity image rendering, we propose to mitigate the issue by initializing and fine-tuning the pre-trained model from a seed image. The images are then progressively registered and added to the training buffer, which is used to train the model further. We also propose to refine the camera poses and pointmaps by minimizing a point-to-camera ray consistency loss across multiple views. When evaluated on diverse challenging datasets, our method outperforms state-of-the-art pose-free NeRF/3DGS methods in terms of both camera pose
Yu Chen, Rolandos Alexandros Potamias, Evangelos Ververas et al.· Neural Information Processin...· 0 citations
Multi-Camera People Tracking (MCPT) traditionally relies on precise intrinsic and extrinsic camera calibration to project 2D detections into a unified 3D world coordinate system.However, manual calibration constitutes a major bottleneck in large-scale dataset generation from unconstrained video archives. This work proposes a unified calibration-free 3D MCPT framework that infers geometric structure directly from visual data using deep foundation models. The system integrates anchor-free detection (YOLOX), robust tracking (BoT-SORT), omni-scale appearance embedding (OsNet), pose estimation (HRNet via MMPose), and transformer-based geometric reconstruction using the Visual Geometry Grounded Transformer (VGGT). A pose-guided 3D lifting strategy projects head keypoints onto a reconstructed manifold, eliminating dependence on ground-plane homography. Global identity association is formulated as hierarchical agglomerative clustering under a joint appearance-geometry cost with strict velocity gating. Evaluation on the AI City Challenge 2024 demonstrates a HOTA score of 53.13% without access to ground-truth calibration matrices, establishing a strong baseline for purely vision-based 3D tracking.
Immersive digital experiences rely increasingly on high-quality 3D content, yet traditional 3D authoring demands specialized expertise, multi-camera capture rigs, and prolonged processing pipelines that remain out of reach for most developers. This paper presents a lightweight, end-to-end artificial intelligence framework that automatically reconstructs a textured 3D polygon model from a single RGB photograph and renders it interactively through a browser-native WebAR interface. The system integrates a deep learning model utilizing convolutional layers to capture representative patterns and features. with a monocular depth-estimation module, Marching-Cubes mesh synthesis, UV-texture projection, and progressive mesh-simplification optimized for delivery over A-Frame, Three.js, and the WebXR Device API. Empirical evaluation confirms a reconstruction accuracy of 96.8% across benchmark image sets, with processing times within practical thresholds on commodity mobile hardware. The proposed architecture eliminates dedicated application installation and offers a scalable, cost-efficient route to AI-powered 3D asset generation.
Shaik Farheen Taj, Ramesh Shahabadkar· International Research Journ...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.