A georeferenced NeRF-based UAV acquisition framework for automated waypoint planning and supervised close-proximity execution and demonstrates improved close-range flight proximity, photographic accuracy, and 3D reconstruction fidelity compared with the evaluated baselines.
Abstract
High-precision 3D reconstruction of objects with complex surfaces, such as ancient architecture and detailed artworks, requires close-range image acquisition, which remains challenging for Unmanned Aerial Vehicle (UAV) systems. The operational proximity of current UAV workflows is often insufficient to capture fine geometric and textural details, limiting high-fidelity digitization. This paper presents a georeferenced NeRF-based UAV acquisition framework for automated waypoint planning and supervised close-proximity execution. The core of the framework is a path-planning module that operates on a metric geometric prior established through Geographic Neural Radiance Fields (Geo-NeRF), which denotes a georeferenced NeRF modeling pipeline rather than a new NeRF architecture or loss function. By generating waypoints directly on this neural representation and optimizing the flight path via a nearest-neighbor strategy, the proposed framework supports close-proximity image acquisition for static targets under controlled conditions. Empirical validation demonstrates improved close-range flight proximity, photographic accuracy, and 3D reconstruction fidelity compared with the evaluated baselines.
Abstract. High-resolution 3D documentation of cultural heritage sites is essential for their preservation. While terrestrial laser scanning (TLS) remains the gold standard, it is often cost-intensive compared to photogrammetry. This study evaluates three image-based reconstruction techniques, Multi-View Stereo (MVS), Neural Radiance Fields (NeRF) and 3D Gaussian Splatting (3DGS), by applying them to a complex scene featuring a chapel and its surrounding vegetation, sensed from an uncrewed aerial vehicle (UAV). A hybrid TLS/MVS model provides a high-accuracy reference. Using identical interior and exterior camera parameters of the 105 UAV-acquired images, we generate dense point clouds with all methods and assess geometric accuracy and completeness using the M3C2 algorithm. Results show that MVS achieves superior accuracy (standard deviation of all M3C2 distances: MVS = 0.11 m, NeRF = 0.15 m), whereas NeRF attains up to 20% higher completeness, particularly in low-texture and vegetation-occluded regions. The 3DGS point cloud was deemed too sparse and was therefore not used for further analysis. The study highlights the potential of NeRFs to recover partially occluded or sparsely textured geometries that are challenging for MVS and suggests a complementary use of both approaches for cost-efficient documentation of cultural heritage.
Frederik Schulte, P. Akwensi, L. Winiwarter· The International Archives o...· 0 citations
Abstract. The proliferation of continuous Neural Radiance Field (NeRF) and 3D Gaussian Splatting (3DGS) has shifted the paradigm of 3D aerial reconstruction from relying solely on geometric stereo matching to inverse rendering optimization. However, while these emerging rendering-based frameworks excel in synthesizing photo-realistic novel views, their capability to extract accurate surfaces in complex aerial scenarios remains ambiguous compared to traditional methods. To establish a clearer understanding, this study presents a comprehensive evaluation of five representative frameworks spanning traditional Structure from Motion (SfM), purely Signed Distance Field (SDF) representations, unstructured 3D Gaussians, hybrid voxel-Gaussians, and strictly explicit sparse voxels. By systematically standardizing identical computational environments, inputs, and unified mesh-extraction pipelines on both real-world airborne LiDAR datasets and synthetic cityscapes, we assess their performance regarding visual fidelity, geometric accuracy, and resource efficiency. The experimental results reveal that while traditional MVS produces the highest overall geometric precision by strictly enforcing multi-view rigid geometry, it is prone to failures in texture-less regions. Among rendering-based representations, a fundamental trade-off exists: highly flexible, unstructured 3DGS achieve highest visual scores but degrade the underlying geometric surfaces; conversely, explicitly structured techniques demonstrate distinct superiority in regularizing topological coherence and floating artifact suppression. Furthermore, we observe that integrating structured voxels avoids the severe memory bottlenecks associated with extracting geometries from chaotic unorganized primitives. These empirical findings emphasize that for large-scale aerial photogrammetry, integrating explicit spatial structuralization into differentiable rendering pipelines is imperative for achieving scalable operations and bridging the geometric accuracy gap with traditional methods.
Shihan Chen, Zhaojin Li, Q. Yan et al.· The International Archives o...· 0 citations
Accurate spatial localization of small, transient targets in low-texture aquatic environments remains a fundamental challenge in UAV-based remote sensing, where open-water surfaces often lack stable tie points, degrading exterior orientation estimation and conventional photogrammetric georeferencing. An integrated UAV framework combining DG/AAT-BA georeferencing with deep-learning-based oriented bounding box (OBB) detection was implemented for high-precision localization, validated on the Critically Endangered Yangtze finless porpoise (YFP, Neophocaena asiaeorientalis) in the Yangtze–Poyang Lake system. The georeferencing component selects direct georeferencing (DG) in open-water scenes and automated aerial triangulation with bundle adjustment (AAT-BA) in feature-rich nearshore scenes. Validation using two static verification points showed that, relative to DG, AAT-BA reduced geometric georeferencing RMSE from 2.59 to 0.62 m under straight-flight conditions and from 3.61 to 0.67 m under turning-flight conditions. For target detection, a lightweight Laplacian edge-enhancement convolution module (LapConv) was incorporated into YOLO-OBB backbones, amplifying weak-edge and low-contrast features of partially submerged targets. Across four representative YOLO-OBB models and three group-constrained partitions, LapConv consistently improved the mean mAP@0.5, with gains of 0.026, 0.024, 0.019, and 0.026 for YOLOv8, YOLO11, YOLO12, and YOLO26, respectively. Applying this framework to six UAV missions across three ecologically and hydrologically distinct subregions enabled georeferenced mapping of porpoise distributions and visualized spatial distribution characteristics during the survey period. The approach is reproducible, minimally invasive, and potentially transferable to UAV-based monitoring of other small aquatic wildlife, providing a methodological basis for fine-scale spatial surveys and subsequent habitat analysis.
Dongxu Yang, Wanbing Ren, Yanren Li et al.· Drones· 0 citations
Abstract. The 3D documentation of complex scenes—characterized by restricted spaces, irregular geometries, and poor lighting—remains a significant challenge in cultural heritage. This study proposes a rapid data acquisition methodology based on the multi-sensor fusion of Terrestrial Laser Scanning (TLS) and Spherical Photogrammetry (SP). The approach was validated in two distinct complex environments: an ancient Egyptian rock-cut tomb (QH36, Aswan, Egypt) and a natural Iberian sanctuary cave (Cueva de la Lobera, Jaén, Spain). The methodology uses TLS to establish a high-precision geometric backbone, achieving registration errors below 0.5 cm. By extracting Ground Control Points (GCPs) directly from the TLS point cloud, the reliance on traditional total station surveying was significantly reduced, enhancing fieldwork efficiency. SP was implemented to obtain realistic textures and to support geometry by using a 360-degree multi-camera with integrated LED lighting, providing full spherical coverage and high-resolution textures. Results indicate that SP is at least six times faster than conventional photogrammetry. Furthermore, the use of TLS-derived meshes enabled advanced digital masking to remove non-interest objects (e.g., archaeological equipment) from the final models. While conventional photogrammetry remains the benchmark for fine architectural details, this research demonstrates that the TLS-SP fusion is the most viable solution for the rapid, high-accuracy documentation of constrained heritage sites. This hybrid workflow ensures geometric integrity while drastically reducing acquisition times, providing a robust framework for future archaeological and conservation projects.
A. Mozas-Calvache, José Luis Pérez-García, J. M. Gómez-López et al.· The International Archives o...· 0 citations
Abstract. Satellite imagery offers a distinct advantage in Earth observation by providing expansive coverage and enabling the monitoring of inaccessible regions without physical on-site intervention, serving as a significantly more cost-effective and scalable alternative to traditional aerial or ground-based surveys. The task of 3D reconstruction from multi-view satellite images has therefore been a pivotal point of research at the intersection of photogrammetry and remote sensing. Recently, novel-view synthesis techniques such as Neural Radiance Fields (NeRF) and 3D Gaussian Splatting (3DGS) have accelerated the accuracy and speed of topographic modeling. Among these, Earth Observation Gaussian Splatting (EOGS) has emerged as a state-of-the-art approach by adapting 3DGS to handle the unique geometric and radiometric characteristics of satellite data, including Rational Polynomial Coefficients (RPCs) and varying solar conditions. However, the standard EOGS pipeline relies on stochastic initialization, where Gaussians are distributed uniformly within a volumetric bounding box, leading to high computational overhead and dependency on aggressive pruning that can inadvertently remove critical geometric features, particularly in areas with complex urban structures. To address these limitations, we propose Bundle-Adjusted Initialization for Earth Observation Gaussian Splatting, which leverages sparse point clouds from bundle adjustment as geometric priors for Gaussian initialization. Combined with an adaptive densification strategy, our method achieves faster convergence and improved DSM accuracy on the DFC2019 dataset compared to the EOGS baseline.
Jiyong Kim, Shuang Song, Rongjun Qin· The International Archives o...· 0 citations
In this paper, we evaluated image quality, algorithms and a workflow associated with matching horizons derived from 3D terrain data to ground imagery as a way to visually estimate a geographic position. The evaluation represented a passive, vision-based navigation technique using long-standoff (>1 KM) terrain features and a horizon detection algorithm hosted within an Android-based geospatial application. Our method involved quantitatively grading edge image feature quality based on pixel data between sky and terrain from high to poor. We tested the algorithm and processing to derive a position using images of varied quality, representing fine and gross, and near and far geographic terrain structures. The site chosen for our tests was located near the Organ Mountains in New Mexico to take advantage of largely unobstructed, long-distance features that challenged both image quality and horizon detection. Testing used the Samsung S23 Ultra (S23U) phone’s primary internal camera to acquire the necessary ground images and native compute power. Our evaluation workflow featured both pre-processing and near-real-time processing elements for position estimations. Pre-processing involved building a Geopackage containing geolocated, synthetic horizons extracted from available 3D terrain data of the test area and camera/sensor configuration data. These data were pre-loaded onto the phone to accomplish the live, near-real-time positional determinations matched to the extracted horizons generated from images acquired by the S23U camera. Our results showed that single-image processing, where only one high-quality ground photo was acquired, 75% of solutions were within 100 m of the actual camera position (compared with the internal sensor-based, Exchangeable Image File Format (EXIF) metadata). Single images of fair quality resulted in positional accuracies where only 38% of the solutions were within 100 m of the EXIF. Improvement was realized when four or more images from varying directions collected from a single location resulted in over 90% of positions falling within 100 m of the EXIF.
J. Ruby, Jimmy R. Carter, Melissa Pham et al.· Applied Sciences· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.