Skip to content
Open access

Monocular Depth Estimation from UAV Images for 3D Documentation of Architectural Heritage: A Depth Anything V2-Based Approach

Aug 2026 · The International Archives of the Photogrammetry, Remote Sensing and Spatial Information Sciences · 0 citations · 5 references

TL;DR

Although accuracy is still insufficient for demanding metric applications, the results support the use of MDE as a complementary source for thematic interpretation, scene understanding, robotics, navigation, and related tasks where strict geometric precision is not required.

Abstract

Abstract. Monocular depth estimation (MDE) has reached notable maturity in computer vision, yet its application to UAV-based architectural heritage documentation remains underexplored. This study assesses whether the depth foundation model Depth Anything V2 can be transferred from terrestrial to aerial imagery. The analysis relies on MDE4BH, a benchmark of over 3,000 UAV images covering ten heterogeneous heritage scenarios (urban areas, façades, towers, villas, domes, and archaeological sites). Masked photogrammetric depth maps serve as metric reference for calibration, validation, and supervised retraining. Two baseline configurations are evaluated: a relative model with scene-specific linear rescaling and the direct application of the metric model. The rescaled relative model shows acceptable performance in several subsets, whereas the metric model exhibits systematic bias, weak consistency, and scale collapse due to domain shift between terrestrial training data and aerial acquisition geometry. To address these limitations, a two-step fine-tuning strategy is introduced, focusing on the decoder and regression head. The first stage uses mainly oblique UAV images; the second integrates oblique and nadir views to improve viewpoint generalization. The adapted model significantly reduces bias and enhances metric stability across the benchmark. However, residual errors remain spatially structured, with clustering and recurrent artefacts near object boundaries, multi-level roofs, and radiometrically heterogeneous surfaces. Although accuracy is still insufficient for demanding metric applications, the results support the use of MDE as a complementary source for thematic interpretation, scene understanding, robotics, navigation, and related tasks where strict geometric precision is not required.

Read PDF

Similar papers

Open access Jul 2026

Using NeRFs for UAV-based 3D reconstruction of complex scenes: A comparison to MVS

Abstract. High-resolution 3D documentation of cultural heritage sites is essential for their preservation. While terrestrial laser scanning (TLS) remains the gold standard, it is often cost-intensive compared to photogrammetry. This study evaluates three image-based reconstruction techniques, Multi-View Stereo (MVS), Neural Radiance Fields (NeRF) and 3D Gaussian Splatting (3DGS), by applying them to a complex scene featuring a chapel and its surrounding vegetation, sensed from an uncrewed aerial vehicle (UAV). A hybrid TLS/MVS model provides a high-accuracy reference. Using identical interior and exterior camera parameters of the 105 UAV-acquired images, we generate dense point clouds with all methods and assess geometric accuracy and completeness using the M3C2 algorithm. Results show that MVS achieves superior accuracy (standard deviation of all M3C2 distances: MVS = 0.11 m, NeRF = 0.15 m), whereas NeRF attains up to 20% higher completeness, particularly in low-texture and vegetation-occluded regions. The 3DGS point cloud was deemed too sparse and was therefore not used for further analysis. The study highlights the potential of NeRFs to recover partially occluded or sparsely textured geometries that are challenging for MVS and suggests a complementary use of both approaches for cost-efficient documentation of cultural heritage.

Frederik Schulte, P. Akwensi, L. Winiwarter · 0 citations
Jul 2026

DAPM: UAV Monocular Depth Estimation from Any Height, Pitch, Roll and FOV

Monocular depth estimation is a fundamental prerequisite for 3D reconstruction and autonomous navigation in Unmanned Aerial Vehicles (UAVs). In practical deployments, UAVs operate under highly dynamic camera poses characterized by continuous variations in height, pitch, roll, and field of view (FOV). Existing monocular depth estimation methods frequently fail to generalize across such diverse perspectives and the expansive scale of depth distributions inherent in aerial scenes. To address these challenges, we establish a quantitative representation of UAV viewing angles through rigorous theoretical analysis, deriving the geometric correspondence between viewing angles and view distances using the ground plane as a reference for observation. Building upon this, we propose Depth Estimation for Any Perspectives Model (DAPM), representing the first monocular framework specifically designed for UAV aerial imagery to jointly estimate camera pose and depth under continuously varying viewpoints. Specifically, we introduce an Ideal Ground Depth (IGD) module that leverages the derived geometric relationships between UAV perspectives and view distances to implement dense camera-pose supervision and enhance depth features. And we further develop a coarse-to-fine Progressive Quantization Bins (PQB) module. By incorporating progressive supervision and hierarchical quantization bins, the PQB module enables robust estimation in complex UAV aerial imagery. To evaluate the proposed framework, we present the UAV Any Perspectives Depth (UAPD) dataset, featuring comprehensive and continuous distributions of pose parameters. Experimental results on UAPD demonstrate that DAPM achieves state-of-the-art performance across both depth and camera-pose estimation metrics. The source code and datasets are available at: https://github.com/ThisIsLT/DAPM.

Tong Ling, Wenhui Diao, Yingchao Feng et al. · 0 citations
Preprint Aug 2026

Iterative Hybrid Discrete-Continuous Viewpoint Planning for UAV Photogrammetry

Unmanned aerial vehicle (UAV) photogrammetry requires camera networks that provide sufficient surface coverage, image overlap, parallax, and resolution, yet conventional flight patterns are often poorly adapted to scene geometry resulting in local reconstruction errors. This paper proposes an iterative hybrid discrete-continuous viewpoint planning method for targeted UAV photogrammetry from a proxy reconstruction. The method scores sampled surface points using photogrammetric heuristics based on frontality, imaging distance, parallax, and multi-view observation count, while also evaluating the full viewpoint set in terms of visibility, pairwise overlap, and graph connectivity. Candidate viewpoints are generated around weakly observed regions, refined using clustered Covariance matrix adaptation evolution strategy (CMA-ES) optimisation, and removed when redundant. The final flight path combines close-range detail viewpoints with wider model-coverage viewpoints, balancing local reconstruction quality with global image-network robustness. Evaluation on three synthetic scenes shows that the proposed method improves both reconstruction accuracy and completeness compared with prior UAV path-planning methods.

Alana Grech, Daniel Pisani, Andrew Grima et al. · 0 citations
Preprint Aug 2026

SiZeUp: Fast 3D Proxy from Aerial Images via Depth Ordinal Loss

We present SiZeUp, a fast and scalable approach for constructing large-scale 3D urban proxy models directly from calibrated oblique aerial imagery. Our method adopts a height-from-footprint representation, reducing 3D building abstraction to a low-dimensional optimization problem in which building footprints are extruded by a single height parameter. To enable efficient and robust height estimation, we introduce an ordinal depth consistency loss that enforces agreement between the relative depth ordering of rendered proxies and depth priors predicted by a monocular depth model. This is realized through a differentiable renderer that maps parametric building proxies into multi-view depth images, allowing gradients to be propagated from depth supervision to building heights. Our ordinal formulation produces stable optimization in practice and avoids explicit feature matching or dense point cloud reconstruction. Rather than relying on metric depth, which can be unreliable under monocular scale ambiguity, our ordinal depth consistency loss operates on relative depths, providing a more reliable signal across views. Combined with an efficient dynamic view selection, our approach achieves a 23-52$\times$ speedup over state-of-the-art proxy reconstruction pipelines while maintaining comparable proxy-level coverage and volume consistency, making it well suited for large-scale urban modeling tasks.

Wenjun Zhou, Yunshan Li, Qiaoyu Zhu et al. · 0 citations
Review Open access Jul 2026

AI-Driven 3D reconstruction and quality assessment for Cultural Heritage: first results from the HERITALISE project

The first results of the AI-based processing pipeline developed within the HERITALISE project are presented, applied to three multiscale case studies at the Reggia di Venaria Reale, demonstrating strong photorealistic rendering capabilities, particularly for complex material properties and geometrically challenging interiors, whilst highlighting current limitations for metric surveying applications.

F. Chiabrando, A. Lingua, Alessio Martino et al. · 0 citations
Review Open access Jul 2026

Fast acquisition for modelling heritage-related complex scenes based on TLS and spherical photogrammetry

Abstract. The 3D documentation of complex scenes—characterized by restricted spaces, irregular geometries, and poor lighting—remains a significant challenge in cultural heritage. This study proposes a rapid data acquisition methodology based on the multi-sensor fusion of Terrestrial Laser Scanning (TLS) and Spherical Photogrammetry (SP). The approach was validated in two distinct complex environments: an ancient Egyptian rock-cut tomb (QH36, Aswan, Egypt) and a natural Iberian sanctuary cave (Cueva de la Lobera, Jaén, Spain). The methodology uses TLS to establish a high-precision geometric backbone, achieving registration errors below 0.5 cm. By extracting Ground Control Points (GCPs) directly from the TLS point cloud, the reliance on traditional total station surveying was significantly reduced, enhancing fieldwork efficiency. SP was implemented to obtain realistic textures and to support geometry by using a 360-degree multi-camera with integrated LED lighting, providing full spherical coverage and high-resolution textures. Results indicate that SP is at least six times faster than conventional photogrammetry. Furthermore, the use of TLS-derived meshes enabled advanced digital masking to remove non-interest objects (e.g., archaeological equipment) from the final models. While conventional photogrammetry remains the benchmark for fine architectural details, this research demonstrates that the TLS-SP fusion is the most viable solution for the rapid, high-accuracy documentation of constrained heritage sites. This hybrid workflow ensures geometric integrity while drastically reducing acquisition times, providing a robust framework for future archaeological and conservation projects.

A. Mozas-Calvache, José Luis Pérez-García, J. M. Gómez-López et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.