Global visual localization of unmanned aerial vehicles (UAVs) using remote-sensing reference maps has attracted increasing attention. However, differences in acquisition time and imaging platform between UAV and reference imagery introduce substantial cross-domain appearance and viewpoint shifts, challenging robust six-degree-of-freedom (6-DoF) pose estimation. To mitigate these shifts, we render UAV-viewpoint references from Google 3D Tiles across locations, altitudes, and orientations. A two-stage strategy adapts SALAD with pose-near positives and geographically distant hard negatives; local geometric consistency then re-ranks the Top-K candidates. We further propose Retrieval-In-Matching (RIM), which freezes the adapted DINOv2-B retriever and distills a local-descriptor decoder from its token field and a shallow VGG19 detail stream. One query-side DINOv2-B backbone forward therefore supports both SALAD retrieval and local description, eliminating a second foundation-model backbone while preserving the retrieval descriptors by construction. We evaluate RIM zero-shot on the reconstructed EPFL Urbanscape and self-collected Chang'an Park datasets, both geographically disjoint from the training data. RIM outperforms ten retrieval baselines. Under the full 3D distance metric at 25/50 m, it improves Recall@1 over SALAD by 8.55/13.77 percentage points on EPFL and 4.45/8.94 points on Park. At Top-K=5, the measured online query path through retrieval, candidate matching, and robust geometric verification takes 90.8 ms: 1.2 times faster than the strongest separate sparse-matching baseline and over 30 times faster than RoMa, while maintaining comparable re-ranking accuracy. These results demonstrate an efficient UAV global visual localization pipeline under unreliable satellite navigation. The source code is available at https://github.com/curious-energy/RIM.
Xin Li, Si-Yuan Duan, Shang Wang et al.· arXiv.org· 1 citation
Global visual localization of unmanned aerial vehicles (UAVs) using remote-sensing reference maps has attracted increasing attention. However, acquisition-time and imaging-platform differences between UAV and reference imagery induce substantial cross-domain appearance and viewpoint shifts, challenging robust six-degree-of-freedom (6-DoF) pose estimation. We address these shifts by sampling UAV-viewpoint reference views from Google 3D Tiles across locations, altitudes, and orientations. A two-stage cross-domain fine-tuning recipe adapts SALAD using pose-near positives and geographically distant hard negatives, while local geometric consistency re-ranks the Top-K candidates. We further propose Retrieval-In-Matching (RIM), which freezes the adapted DINOv2-B retriever and distils a local-descriptor decoder that reuses its token field alongside a shallow VGG19 detail stream. One query-side DINOv2-B forward thus serves both SALAD retrieval and local description, eliminating a second foundation-model backbone while preserving retrieval descriptors by construction. We evaluate RIM zero-shot on the reconstructed EPFL Urbanscape and self-collected Chang'an Park datasets, both geographically disjoint from the training data. RIM outperforms ten recent retrieval baseline families. At 25/50 m under the full 3D distance metric, it improves Recall@1 over SALAD by 8.55/13.77 percentage points on EPFL and 4.45/8.94 points on Park. At Top-K=5, the complete measured localization query, including retrieval, candidate matching, and robust geometric verification, takes 67.9 ms end-to-end: 1.8 times faster than the strongest separate sparse-matching baseline and over 40 times faster than RoMa, while achieving comparable re-ranking accuracy. These results establish an efficient and deployable pipeline for UAV global visual localization in GNSS-challenged environments.
Xin Li, Siyuan Duan, Shang Wang et al.· arXiv.org· 1 citation
Multiview geo-localization utilizing drone and satellite imagery offers a reliable alternative to GPS-based positioning in challenging environments such as urban canyons and electromagnetically degraded areas. A key challenge in this task stems from the spatial misalignment between drone-view and satellite-view images, particularly when target buildings appear off-center due to varying flight attitudes and environmental disturbances. Existing methods, which often rely on implicit center-aligned assumptions, show limited robustness under such offset conditions. To address this limitation, we construct a novel drone-view offset dataset named Offset-1652 by applying controlled translational transformations to the benchmark dataset, simulating realistic displacement scenarios along horizontal, vertical, and diagonal directions. Furthermore, we propose a large-kernel perceptual attention network (LK-PAN) that employs large-kernel depthwise convolutions to expand the receptive fields, thereby capturing global contextual information even for off-center targets. A symmetric InfoNCE loss is introduced to enhance cross-modal feature alignment and improve discrimination of hard negative samples. Comprehensive experiments conducted on both established benchmark datasets and the newly developed Offset-1652 and Offset-1652-MIX datasets demonstrate that the proposed method significantly outperforms existing approaches under a wide range of offset conditions. These results confirm its superior robustness and generalization capability for practical multiview geo-localization in complex operational environments. The source code and datasets are available at https://github.com/HAORANJY/LK-PAN-main
Bangyong Sun, Mian Li, Weifeng Wang et al.· IEEE Transactions on Geoscie...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.