AirAlign is proposed, a framework for RGB-only image-pair relative pose alignment for UAVs, using a pretrained visual geometry reconstruction model as the backbone to extract geometry-aware features from source-target image pairs.
Abstract
Unmanned aerial vehicle (UAV) navigation in modern low-altitude environments requires more accurate pose alignment in the final approach stage for target information acquisition or manipulation, making"last-meter"navigation increasingly important. However, severe viewpoint and appearance variations make this task challenging. To tackle this problem, we propose AirAlign, a framework for RGB-only image-pair relative pose alignment for UAVs. AirAlign uses a pretrained visual geometry reconstruction model as the backbone to extract geometry-aware features from source-target image pairs. In addition, to better utilize the limited training data, we split the training set into multiple scene-disjoint folds for unseen cross-validation and model selection. During inference, the predictions of the selected models are averaged to form the ensemble output of the overall framework. Experiments on the PairUAV challenge at the ACMMM 2026 Workshop on UAVs in Multimedia demonstrate the effectiveness and robustness of our method, while comprehensive ablation studies validate the contribution of each component.
Abstract. Reliable visual localization is essential for long-range VTOL UAV mapping in GNSS-degraded environments. This paper presents a quantitative evaluation framework for monocular ORB-SLAM3 using a 66.48 km multi-altitude UAV mission and aerial-triangulation-derived camera poses as reference data. The workflow associates SLAM and reference trajectories by image key, applies Sim(3)-based metric alignment, corrects coordinate-axis inconsistency, and refines attitude by a global rotation offset, enabling full-mission and segment-level comparison in a common metric frame. The evaluation covers four altitude segments, namely 100, 150, 200, and 250 m AGL, under three protocols: No-Loop (NL), With-Loop Global Slice (GS), and With-Loop Local Re-Sim(3) (LR). For the full mission, the proposed alignment achieves a 3D position RMSE of 7.41 m over 5330 matched frames and substantially reduces the geometric deformation observed in the S+T baseline. Segment-level results show a strong altitude dependency in the isolated NL runs, with 3D RMSE decreasing from 22.95 m at 100 m to 5.49 m at 250 m. Among the three protocols, LR consistently yields the best segment-level position accuracy, reaching 4.00, 8.26, 3.94, and 3.92 m at 100, 150, 200, and 250 m, respectively. Long-range analysis further shows that the trajectory remains globally bounded, while cumulative 3D endpoint drift increases from 0.35 m at 50 m to 10.66 m at 25.6 km. These results indicate that ORB-SLAM3 can support large-scale trajectory estimation for UAV mapping, but its evaluated quality depends strongly on alignment, segmentation, and evaluation strategy.
Ming-Jyun Yang, J. Jhan, Runmeng Tang· The International Archives o...· 1 citation
Unmanned aerial vehicles (UAVs) increasingly require robust visual localization in GNSS-denied environments. A common solution estimates UAV poses by matching keypoints between UAV images and geo-tagged orthographic reference maps derived from satellite or aerial imagery, followed by Perspective-\(n\)-Point (PnP) pose solving. However, such reference maps mainly record top-down surfaces such as roofs and ground planes, while vertical structures such as facades and walls are often compressed or missing. Consequently, many visually distinctive keypoints in low-altitude UAV images have no valid counterparts in the reference map, leading to redundant matches and inaccurate pose estimation. To address this issue, we propose DECO, a DEpth-guided CO-visibility reasoning framework for low-altitude UAV visual localization. DECO uses monocular depth priors to infer local surface geometry and estimate co-visible regions between UAV images and the reference map. Based on this prior, a Geometry-Saliency Coupled Co-visibility Score is introduced to jointly consider geometric co-visibility and detector saliency for keypoint ranking. In this way, DECO retains keypoints that are both visually distinctive and geometrically co-visible, improving feature matching and PnP-based pose estimation. Extensive experiments demonstrate that DECO achieves superior localization performance and can be integrated with different depth models, feature detectors, and matchers. The source code will be available at https://github.com/UAV-AVL/DECO.
Yi-Bin Ye, Xichao Teng, Shuo Chen et al.· 0 citations
Global visual localization of unmanned aerial vehicles (UAVs) using remote-sensing reference maps has attracted increasing attention. However, acquisition-time and imaging-platform differences between UAV and reference imagery induce substantial cross-domain appearance and viewpoint shifts, challenging robust six-degree-of-freedom (6-DoF) pose estimation. We address these shifts by sampling UAV-viewpoint reference views from Google 3D Tiles across locations, altitudes, and orientations. A two-stage cross-domain fine-tuning recipe adapts SALAD using pose-near positives and geographically distant hard negatives, while local geometric consistency re-ranks the Top-K candidates. We further propose Retrieval-In-Matching (RIM), which freezes the adapted DINOv2-B retriever and distils a local-descriptor decoder that reuses its token field alongside a shallow VGG19 detail stream. One query-side DINOv2-B forward thus serves both SALAD retrieval and local description, eliminating a second foundation-model backbone while preserving retrieval descriptors by construction. We evaluate RIM zero-shot on the reconstructed EPFL Urbanscape and self-collected Chang'an Park datasets, both geographically disjoint from the training data. RIM outperforms ten recent retrieval baseline families. At 25/50 m under the full 3D distance metric, it improves Recall@1 over SALAD by 8.55/13.77 percentage points on EPFL and 4.45/8.94 points on Park. At Top-K=5, the complete measured localization query, including retrieval, candidate matching, and robust geometric verification, takes 67.9 ms end-to-end: 1.8 times faster than the strongest separate sparse-matching baseline and over 40 times faster than RoMa, while achieving comparable re-ranking accuracy. These results establish an efficient and deployable pipeline for UAV global visual localization in GNSS-challenged environments.
Xin Li, Siyuan Duan, Shang Wang et al.· arXiv.org· 1 citation
Unmanned aerial vehicle (UAV) multimodal perception integrates visible (RGB), infrared (IR), synthetic aperture radar (SAR), and depth sensors for scene understanding under diverse conditions. However, differences in optics, resolution, and mounting often limit practical systems to global or image-center alignment. After tokenization, parallax, platform motion, and lens distortion can shift corresponding patch centers across modalities, weakening the spatial correspondence assumed by dense contrastive learning and cross-modal fusion. We propose GAAT (Geometry-Aware Alignment Transformer), an alignment-first pretrained model that estimates local correspondence reliability before cross-modal interaction. GAAT introduces syncPATC, which learns patch-center consistency under synchronized view transformations without correspondence annotations. It emits geometric priors, including token and query confidence, query centers, and sub-token offsets, that identify reliable local anchors across residual misalignment. Guided by these priors, MG-Sparse-MMA performs query-mediated sparse fusion over top-K_s reliable regions, replacing dense all-patch interaction with geometry-calibrated local updates. RA-QCGCL aligns pretraining supervision with this sparse query bottleneck through reliable patch-to-patch, patch-to-query, and query-to-query contrastive branches. We introduce UAVMeta and StateBench, which provide four acquisition-state scores derived from platform telemetry and image statistics: camera reliability, observation scale, viewpoint stability, and flight maneuver complexity. Extensive experiments across six downstream tasks demonstrate consistently superior transfer performance, establishing GAAT as a state-of-the-art multimodal foundation model for UAV perception. StateBench further enables a systematic diagnosis of real-world acquisition conditions.
Jing-Pu Yang, Deming Tang, Yi-Lin Sun et al.· 1 citation
Abstract. Navigating Unmanned Aerial Vehicles (UAVs) in Global Navigation Satellite System (GNSS)-denied environments requires reliable autonomous localization techniques. This study proposes a vision-based localization framework utilizing satellite true orthophotos and Digital Surface Models (DSMs) as absolute geospatial references. The algorithmic pipeline integrates deep learning architectures—specifically SuperPoint and LightGlue—to establish robust image-to-map feature correspondences. The matched correspondences are used to estimate camera exterior orientation parameters through collinearity-based spatial resection with an Iteratively Reweighted Least Squares (IRLS) approach. To validate the proposed methodology, a multi-altitude dataset (100–250 m) was acquired across structurally diverse terrains, including dense building, high vegetation, and bare ground areas. Experimental evaluations demonstrate that the framework achieves meter-level absolute positioning accuracy and stable pose estimation. Analyses further reveal that matching robustness and localization success rates depend heavily on terrain texture and flight altitude; geometrically structured urban scenes at moderate-to-high altitudes consistently yield reliable correspondences, whereas low-texture environments and lower flight altitudes present persistent challenges for continuous visual tracking.
Tai-Cyuan Wang, Lai-Han Tsou, J. Jhan et al.· The International Archives o...· 0 citations
AeroMotion6D is proposed, a temporal transformer-based framework for monocular UAV 6D pose estimation from RGB video that consists of an adaptive context fusion mechanism that can incorporate past context information into the current estimation process and a persistent pose memory module that can convey pose-related information in two consecutive frames.
Mohammad Al Qaderi, M. Hayajneh, Alaa Alghazo et al.· Robotics· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.