Skip to content

DGSfM: Depth-Guided Scale-Aware Global Structure-from-Motion

Jul 2026 · arXiv.org · Vol abs/2607.09507 · 0 citations · 69 references
Computer Science

TL;DR

DGSfM is proposed, a depth-aware global SfM pipeline that uses monocular depth maps as a scalable prior while preserving explicit multi-view optimization, and consistently improves over strong global SfM baselines across sparse and dense matching front-ends, achieving substantial gains in pose accuracy.

Abstract

Global Structure-from-Motion (SfM) is an efficient paradigm for recovering camera poses and sparse 3D structure from unordered images. However, its reliance on scale-ambiguous epipolar geometry makes global positioning sensitive to noisy baseline estimates and weak view-graph constraints, while false edges from visually ambiguous pairs can further degrade reconstruction. We propose DGSfM, a depth-aware global SfM pipeline that uses monocular depth maps as a scalable prior while preserving explicit multi-view optimization. For each image pair, we use a depth-aware relative pose solver to convert scale-ambiguous epipolar constraints into scale-aware relative pose constraints. We further improve robustness through view-graph filtering and depth-consistency-based correspondence pruning, which suppress false edges and matches that remain plausible under epipolar geometry alone. Finally, global scale averaging and depth-guided pose-point initialization align monocular depth maps into a common reconstruction scale and provide stable initialization for global positioning and bundle adjustment. Experiments on ETH3D and IMC2021 show that DGSfM consistently improves over strong global SfM baselines across sparse and dense matching front-ends, achieving substantial gains in pose accuracy. Code is available at https://github.com/sithu31296/DGSfM.

View source

Similar papers

Preprint Aug 2026

Robust Global Structure-from-Motion via View Graph Pruning

Structure-from-Motion (SfM) aims to estimate camera poses and reconstruct 3D structures from a collection of unordered images. Compared with incremental SfM, global SfM achieves better scalability by jointly estimating camera poses based on a view graph constructed from pairwise correspondences. However, its performanc...

Jiamin Xu, Lixing Yao, Weichen Dai et al. · 0 citations
Jul 2026

VidMap: Exploiting Temporal Structure for Video-Based Structure-from-Motion

This work introduces a system that combines the strong sequential constraints of SLAM with the flexibility and global optimization of offline SfM, enabling the metric reconstruction of arbitrary, long, uncalibrated videos.

Zador Pataki, Paul-Edouard Sarlin, Marc Pollefeys · 1 citation · ⚡1
Preprint Sep 2026

HiSfM: Disambiguating Structure-from-Motion via Scaffold-Anchored Hierarchical Reconstruction

Structure-from-Motion (SfM) is a fundamental tool for sparse 3D reconstruction with broad impact in robotics and vision, supporting mapping, localization, and large-scale scene modeling. However, conventional pipelines often fail under hard visual ambiguity caused by repeated or symmetric structures, and incur heavy co...

Zi-Ding Zhao, Hainan Cui, Pei-Lin Tao et al. · 0 citations
#graph neural networks Preprint Sep 2026

Learning Global Camera Poses from Noisy View-Graphs for Structure from Motion

This work presents a deep, global Structure-from-Motion framework based on learned view-graph aggregation that employs a permutation-equivariant, edge-conditioned graph neural network that takes noisy pairwise relative poses as input and outputs globally consistent camera extrinsics.

Fadi Khatib, M. Galun, R. Basri · 0 citations
Preprint Aug 2026

D^2-4DGS: Dual-Depth Guided Sparse-Camera 4D Gaussian Splatting

This work aligns monocular estimates with valid multi-view geometric depths and verify their consistency to identify reliable geometric anchors, which support consistency-aware pruning and depth supervision, and aligned mono-only estimates and RGB-D joint optimization improves appearance fidelity and geometric consiste...

Jijian Zhao · 0 citations
Preprint Sep 2026

RIDE: Relocalization-Informed Depth Estimation with 3D Gaussian Splatting

Render--match--PnP relocalization establishes correspondences between query image pixels and 3D map points for camera pose recovery, but their potential to support dense depth estimation is often overlooked. To exploit this geometric information, we present RIDE, which estimates dense metric depth from a robot's RGB st...

Jia-Rong Lian, Zhen-Hua Xiao, Zhao-Yang Zhang et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.