Aug 2026· PFG – Journal of Photogrammetry Remote Sensing and Geoinformation Science· 0 citations· 234 references
TL;DR
This work aims to guide future research toward robust, accurate, and certifiable 3D reconstruction systems suitable for engineering, industrial, and geospatial applications, by bridging the gap between classical photogrammetry and data-driven 3D vision.
Abstract
Image-based 3D reconstruction is vital in many applications, such as digital twins, smart cities, machine vision, and autonomous driving. In recent years, it has undergone a paradigm shift, propelled by advancements in both conventional photogrammetry and deep learning. This review provides a comprehensive photogrammetric perspective on both conventional and learning-based techniques, a viewpoint that prioritizes geometric fidelity, robustness, handling of uncertainty, and suitability for real-world applications. We first systematically revisit the fundamentals of traditional pipelines: Structure from Motion (SfM), Multi-View Stereo (MVS), and surface reconstruction. The review then details recent progress in conventional methods, highlighting innovations in scalable and efficient SfM, specialized camera models for MVS, and robust surface reconstruction algorithms. Subsequently, we explore the transformative evolution brought by learning-based techniques, including deep SfM, learning-based MVS, differentiable rendering-based scene representation methods (NeRF, 3DGS), groundbreaking feed-forward 3D reconstruction models (e.g., DUSt3R, VGGT), and surface reconstruction including explicit and implicit methods. Emphasis is placed on evaluating whether learning-based approaches genuinely meet photogrammetric requirements such as metric accuracy and reliability, rather than optimizing solely for visual realism.
Furthermore, we conclude by identifying key challenges and research frontiers including generalization across domains, scalability to high-resolution imagery, real-time performance, and uncertainty quantification. By bridging the gap between classical photogrammetry and data-driven 3D vision, this work aims to guide future research toward robust, accurate, and certifiable 3D reconstruction systems suitable for engineering, industrial, and geospatial applications.
Abstract. Monocular depth estimation (MDE) infers depth from a single image, offering significant advantages in computational efficiency and memory consumption compared to conventional Multi-View Stereo (MVS) methods. However, most MDE methods suffer from poor multi-view geometric consistency, which limits their application to photogrammetric 3D reconstruction. To address this issue, this paper employs sparse point clouds of Structure-from-Motion (SfM) as extra geometric constraints and proposes a framework that achieves photogrammetric 3D reconstruction using off-the-shelf learning-based MDE models without the need for additional fine-tuning. Specifically, when SfM priors are available during inference, globally geometrically consistent depth maps can be directly predicted. Otherwise, the estimated monocular depths are aligned to a consistent scale using SfM results via a post-correction step. The resulting depth maps are then fused using a truncated signed distance function (TSDF) to generate dense 3D reconstructions. Experiments on photogrammetric datasets demonstrate that the proposed framework effectively improves geometric consistency across depth maps and enables high-quality scene reconstruction. In addition, we systematically analyze the impact of key parameters in depth inference and fusion, including depth map resolution, voxel size, denoising steps, and ensemble size, on reconstruction performance, and further explore the potential of MDE for photogrammetric 3D reconstruction.
Chunyu Dou, Yifei Yu, Xin Wang et al.· The International Archives o...· 0 citations
This work proposes a hybrid reconstruction pipeline, leveraging the strengths and benefits of each technique, which exploits the accurate geometry of photogrammetry in well-textured regions and the GS capabilities to improve completeness and visual aspect in areas featuring non-collaborative surfaces.
Fabio Remondino, E. M. Farella, Gianluca Bertolasi et al.· The International Archives o...· 0 citations
Recent advancements in transformer-based deep learning have significantly improved the accuracy and efficiency of single-image 3D reconstruction. The transformer-based feed-forward architecture for single-image 3D reconstruction, system integrates DINOv1 vision transformers, triplane repre-sentations and Neural Radiance Fields (NeRF) principles to generate textured 3D meshes from single RGB images. Key contributions include efficient triplane decoding, robust pre-processing with background removal (rembg) and edge de-tection, and a Gradio-based web interface enabling real-time deployment on consumer GPUs (6GB VRAM). Preliminary deployment experiments suggest that Image2Mesh is capable of near-real-time inference on high-performance GPUs while demonstrating a favorable speed–accuracy trade-off compared to traditional photogrammetry-based workflows. Image2Mesh provides a practical foundation for AR/VR, gaming, and content creation applications.
Shravan Shetty, Akash Nayak, Anagha Ankolekar et al.· 2026 International Conferenc...· 0 citations
A semantic-guided 3D Gaussian splatting (3DGS) framework tailored to sparse-view industrial reconstruction was introduced, enabling robust reconstruction from limited viewpoints and offers a practical geometric foundation for automated inspection and remote equipment monitoring.
Boyang Li, Tian-Han Gao, Zuan Gu et al.· Visual Computing for Industr...· 0 citations
Monocular depth estimation has long stood as a fundamental challenge in computer vision, enabling a wide range of applications including 3D reconstruction, robotics, autonomous driving, and augmented reality. This survey traces the field's evolution from early learning-based methods to the emergence of transformative foundation models. We begin by framing the problem, distinguishing between relative and metric depth estimation, and highlighting the key challenges that have shaped a decade of research. We then present common problem formulations and introduce the most widely used datasets, covering indoor, outdoor, and synthetic data. Following this, we review major advances prior to the foundation model era, distilling core insights from influential methods that contributed to improvements in accuracy, efficiency, and robustness. The survey then turns to the recent surge of foundation-model-based approaches, categorizing them into discriminative and generative paradigms and emphasizing the critical roles of large-scale pretraining (e.g., DINOv3) and synthetic data. We compare representative models using both quantitative benchmarks and qualitative examples, and discuss natural extensions to video-based depth estimation. Further, to illustrate real-world impact, we highlight the integration of depth estimation into applications such as visual SLAM, content generation, and robot perception. Finally, we outline open challenges and promising research directions as the field advances further into the era of foundation models.
Mu-Xin Liu, Xiaoyang Lyu, Yang-Tian Sun et al.· 0 citations
Abstract. The paper presents a 3D-informed method for generating stitched panoramic images from multi-camera rigs through rigorous reprojection onto a scene model. Unlike conventional and parallax-tolerant stitching, the proposed approach explicitly accounts for camera calibration, relative orientation, and scene geometry, with the aim of reducing parallax effects while preserving metric consistency. The method is tested on confined environments, where non-coincident projection centres make stitching especially critical, and is evaluated with different rig configurations, including systems with both small and large sensor baselines. Two experiments are performed: a stitching-accuracy test against synthetic reference panoramas, and a Structure from Motion (SfM) test comparing rigorous panoramas with raw fisheye processing. Results show that the proposed approach yields geometrically consistent panoramas and substantial gains in processing efficiency, although raw fisheye images still provide the best overall metric performance in the most demanding reconstruction scenarios.
R. Roncella, L. Perfetti· The International Archives o...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.