Skip to content
Open access

Assessing the Reconstruction Potential of 3D Vision Foundation Models for Oblique Photogrammetry

Jul 2026 · ISPRS Annals of the Photogrammetry, Remote Sensing and Spatial Information Sciences · Vol XI-2-2026, pp. 821-830 · 0 citations · 7 references

TL;DR

Benefiting from the powerful zero-shot generalization, 3D vision foundation models can robustly estimate camera parameters and generate dense point clouds under sparse-view and low-overlap conditions, with some rivaling traditional photogrammetry configured with redundant observations.

Abstract

Abstract. 3D vision foundation models, which directly regress 3D geometry from 2D images in an end-to-end manner, have recently attracted growing attention in the computer vision community. However, their potential for oblique 3D reconstruction has not been systematically evaluated. To this end, we establish an automated evaluation pipeline to benchmark these models on oblique imagery. Our experiments reveal that: benefiting from the powerful zero-shot generalization, 3D vision foundation models can robustly estimate camera parameters and generate dense point clouds under sparse-view and low-overlap conditions, with some rivaling traditional photogrammetry configured with redundant observations. Counterintuitively, two-view reasoning foundation models employing explicit PnP-RANSAC for global alignment consistently outperform multi-view reasoning foundation models inferring multi-view relationships via implicit attention mechanism when processing more than 2 views. Notably, incorporating known camera parameters as conditioning inputs, which act as weak supervision rather than rigid geometric constraints, yields only marginal accuracy improvements. Based on ViT architecture, these foundation models face scalability bottlenecks to large-scale and high-resolution oblique imagery, and their prevalent ideal pinhole camera assumption still makes explicit distortion correction an unavoidable preprocessing step.

Read PDF

Similar papers

Open access Jul 2026

Learning-based Monocular Depth Estimation for Photogrammetric 3D Reconstruction

Abstract. Monocular depth estimation (MDE) infers depth from a single image, offering significant advantages in computational efficiency and memory consumption compared to conventional Multi-View Stereo (MVS) methods. However, most MDE methods suffer from poor multi-view geometric consistency, which limits their application to photogrammetric 3D reconstruction. To address this issue, this paper employs sparse point clouds of Structure-from-Motion (SfM) as extra geometric constraints and proposes a framework that achieves photogrammetric 3D reconstruction using off-the-shelf learning-based MDE models without the need for additional fine-tuning. Specifically, when SfM priors are available during inference, globally geometrically consistent depth maps can be directly predicted. Otherwise, the estimated monocular depths are aligned to a consistent scale using SfM results via a post-correction step. The resulting depth maps are then fused using a truncated signed distance function (TSDF) to generate dense 3D reconstructions. Experiments on photogrammetric datasets demonstrate that the proposed framework effectively improves geometric consistency across depth maps and enables high-quality scene reconstruction. In addition, we systematically analyze the impact of key parameters in depth inference and fusion, including depth map resolution, voxel size, denoising steps, and ensemble size, on reconstruction performance, and further explore the potential of MDE for photogrammetric 3D reconstruction.

Chunyu Dou, Yifei Yu, Xin Wang et al. · 0 citations
Preprint Sep 2026

VI3: Grounding Pretrained 3D Foundation Models with Inertial Cues

3D foundation models (3DFMs) excel at predicting camera poses and dense depth from multiple views of a scene, showcasing strong zero-shot generalization. However, as metric scale is not observable from monocular images, their absolute scale predictions are typically inaccurate. Inertial measurement units (IMUs), present in most devices, naturally complement monocular cameras by observing scaled motion. We introduce VI3, a model-agnostic framework that metrically anchors a pretrained 3DFM using only IMU readings. VI3 initializes and preintegrates the IMU to obtain a metric motion reference, which is then used to recover the scale of the 3DFM outputs. Our method includes adaptable anchoring strategies tailored to diverse 3DFM architectures. Experiments on synthetic and real aerial datasets demonstrate that VI3 recovers metric scale without ground-truth supervision while preserving geometric consistency, acting as a fine refinement under well-conditioned motion and as a strong prior when motion is less informative.

E. Lozano, Alberto Jaenal, Javier Civera · 0 citations
Jul 2026

Axolotl3D: a Unified Framework for Faithful 3D Shape Completion

Axolotl3D is presented, a multi-modal and occlusion-aware 3D generation model that jointly conditions on images, visibility masks, camera parameters, and a partial point cloud that synthesizes diverse conditioning regimes from large-scale 3D data, enabling robust cross-modal reasoning.

A. Hu, Maria Shugrina · 1 citation
Preprint Aug 2026

Confidence matters: Leveraging Multi-view Geometric Priors for GS-based Reconstruction

This work investigates the integration of geometric priors, in the form of predicted normal and depth maps, into the 3DGS framework to improve the reconstruction quality and reveals that multi-view predictions, as they are done by the recent visual geometry grounded transformer (VGGT), outperform single-view alternatives.

Hongyu Zhou, Zorah Lähner · 0 citations
Open access Aug 2026

Semantic-guided 3D Gaussian splatting for sparse-view reconstruction in industrial digital twins

A semantic-guided 3D Gaussian splatting (3DGS) framework tailored to sparse-view industrial reconstruction was introduced, enabling robust reconstruction from limited viewpoints and offers a practical geometric foundation for automated inspection and remote equipment monitoring.

Boyang Li, Tian-Han Gao, Zuan Gu et al. · 0 citations
Open access Jul 2026

Rigorous Projection for Image Stitching: a 3D-Informed Approach for Accurate Panoramic Photogrammetry

Abstract. The paper presents a 3D-informed method for generating stitched panoramic images from multi-camera rigs through rigorous reprojection onto a scene model. Unlike conventional and parallax-tolerant stitching, the proposed approach explicitly accounts for camera calibration, relative orientation, and scene geometry, with the aim of reducing parallax effects while preserving metric consistency. The method is tested on confined environments, where non-coincident projection centres make stitching especially critical, and is evaluated with different rig configurations, including systems with both small and large sensor baselines. Two experiments are performed: a stitching-accuracy test against synthetic reference panoramas, and a Structure from Motion (SfM) test comparing rigorous panoramas with raw fisheye processing. Results show that the proposed approach yields geometrically consistent panoramas and substantial gains in processing efficiency, although raw fisheye images still provide the best overall metric performance in the most demanding reconstruction scenarios.

R. Roncella, L. Perfetti · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.