Abstract. Reliable visual localization is essential for long-range VTOL UAV mapping in GNSS-degraded environments. This paper presents a quantitative evaluation framework for monocular ORB-SLAM3 using a 66.48 km multi-altitude UAV mission and aerial-triangulation-derived camera poses as reference data. The workflow associates SLAM and reference trajectories by image key, applies Sim(3)-based metric alignment, corrects coordinate-axis inconsistency, and refines attitude by a global rotation offset, enabling full-mission and segment-level comparison in a common metric frame. The evaluation covers four altitude segments, namely 100, 150, 200, and 250 m AGL, under three protocols: No-Loop (NL), With-Loop Global Slice (GS), and With-Loop Local Re-Sim(3) (LR). For the full mission, the proposed alignment achieves a 3D position RMSE of 7.41 m over 5330 matched frames and substantially reduces the geometric deformation observed in the S+T baseline. Segment-level results show a strong altitude dependency in the isolated NL runs, with 3D RMSE decreasing from 22.95 m at 100 m to 5.49 m at 250 m. Among the three protocols, LR consistently yields the best segment-level position accuracy, reaching 4.00, 8.26, 3.94, and 3.92 m at 100, 150, 200, and 250 m, respectively. Long-range analysis further shows that the trajectory remains globally bounded, while cumulative 3D endpoint drift increases from 0.35 m at 50 m to 10.66 m at 25.6 km. These results indicate that ORB-SLAM3 can support large-scale trajectory estimation for UAV mapping, but its evaluated quality depends strongly on alignment, segmentation, and evaluation strategy.
Ming-Jyun Yang, J. Jhan, Runmeng Tang· The International Archives o...· 1 citation
Abstract. Navigating Unmanned Aerial Vehicles (UAVs) in Global Navigation Satellite System (GNSS)-denied environments requires reliable autonomous localization techniques. This study proposes a vision-based localization framework utilizing satellite true orthophotos and Digital Surface Models (DSMs) as absolute geospatial references. The algorithmic pipeline integrates deep learning architectures—specifically SuperPoint and LightGlue—to establish robust image-to-map feature correspondences. The matched correspondences are used to estimate camera exterior orientation parameters through collinearity-based spatial resection with an Iteratively Reweighted Least Squares (IRLS) approach. To validate the proposed methodology, a multi-altitude dataset (100–250 m) was acquired across structurally diverse terrains, including dense building, high vegetation, and bare ground areas. Experimental evaluations demonstrate that the framework achieves meter-level absolute positioning accuracy and stable pose estimation. Analyses further reveal that matching robustness and localization success rates depend heavily on terrain texture and flight altitude; geometrically structured urban scenes at moderate-to-high altitudes consistently yield reliable correspondences, whereas low-texture environments and lower flight altitudes present persistent challenges for continuous visual tracking.
Tai-Cyuan Wang, Lai-Han Tsou, J. Jhan et al.· The International Archives o...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.