Skip to content
Open access

DFSA: Dynamic-Feature Collaborative Optimization and Semantic-Alignment Network for UAV Cross-View Geo-Localization

Aug 2026 · Drones · 0 citations · 60 references

TL;DR

This work proposes a novel CVGL method named dynamic-feature collaborative optimization and semantic-alignment network (DFSA), designed to extract robust feature representations and achieve fine-grained alignment, and introduces a semantic segmentation and alignment module that adaptively partitions images into semantic regions based on feature response distributions.

Abstract

Cross-view geo-localization (CVGL) is a critical technology used in unmanned aerial vehicles (UAVs) and widely applied in navigation and target localization tasks. However, owing to the extreme perspective disparity between UAV oblique views and satellite vertical views, CVGL still involves significant challenges, including geometric distortion caused by viewpoint differences, drastic appearance inconsistencies, and the difficulty in bridging semantic gaps between heterogeneous data. To address these issues, we propose a novel CVGL method named dynamic-feature collaborative optimization and semantic-alignment network (DFSA), designed to extract robust feature representations and achieve fine-grained alignment. Specifically, the DFSA employs a residual-based vision transformer as the backbone to capture global context while alleviating the training instability and feature collapse often associated with standard transformers. To bridge the semantic gap between global and local features, we design a feature optimization module comprising a local feature enhancer and a global feature aggregator. This module establishes a closed-loop collaborative system that facilitates top-down semantic guidance and bottom-up detail feedback. Furthermore, we introduce a semantic segmentation and alignment module that adaptively partitions images into semantic regions based on feature response distributions, shifting the matching granularity from the global level to the semantic region level to effectively overcome feature mismatches caused by positional offsets and scale variations. Extensive experiments conducted on the University-1652 and SUES-200 datasets demonstrate the superior image retrieval performance of the proposed DFSA. Specifically, DFSA achieves a Recall@1 of 94.87% and an Average Precision (AP) of 95.32% on the University-1652 dataset and maintains highly competitive Recall@1 performances between 96.83% and 99.25% across various altitudes on the SUES-200 dataset. These results validate the model’s effectiveness in handling extreme viewpoint changes for UAV-based cross-view image retrieval tasks.

Read PDF

Similar papers

Open access Aug 2026

Geo-Consistent Centralized Multi-UAV Gaussian SLAM for Incremental Orthophoto Generation

A centralized GNSS-assisted multi-UAV 3D Gaussian Splatting SLAM framework for online incremental orthophoto mapping that provides a favorable trade-off between geo-consistency, visual fidelity, and efficiency compared with existing methods is presented.

Xiao Zhang, Shuai-Xin Li, Hongbin Dong et al. · 0 citations
Open access Sep 2026

SHALA: Sparse Hierarchical Alignment for Visible-Infrared Vehicle Re-Identification Under Aerial-Ground Viewpoint Gaps

Visible-infrared vehicle re-identification across aerial-ground UAV platforms, where high-altitude and low-altitude views are separated by a large viewpoint gap, supports all-day traffic monitoring and cross-platform target association. However, most existing cross-modal re-identification methods are developed for sing...

Dong Liu, Xiao-Lin Zhao, Si-Yuan Zhao et al. · 0 citations
Preprint Aug 2026

AirAlign: Geometry-Aware Relative Pose Alignment for UAV Last-Meter Navigation

AirAlign is proposed, a framework for RGB-only image-pair relative pose alignment for UAVs, using a pretrained visual geometry reconstruction model as the backbone to extract geometry-aware features from source-target image pairs.

Jin-Yi Zhou, Shuo Feng, Yufei Wu et al. · 0 citations
Aug 2026

Spectrum adaptive enhancement and multi-scale semantic interaction for robust cross-view geo-localization

An adaptive framework for cross-view matching that integrates spectrum adaptive enhancement and multi-scale semantic interaction mechanisms is proposed that leverages frequency-domain information to strengthen structural and detailed features, and employs multi-scale semantic guidance to improve feature alignment.

He Xiao, Zhen-Ju Wen, Yu Du et al. · 0 citations
Open access Sep 2026

Improved 3D Gaussian splatting integrating VGGT and LiDAR: Application in visual relocalization for UAV inspection of large aircraft

For high-precision unmanned aerial vehicle (UAV) visual relocalization in large-scale aircraft inspection, this paper proposes an improved 3DGS method integrating visual geometry grounded transformer (VGGT) with LiDAR and designs a coarse-to-fine two-stage image matching strategy to tackle the challenge. First, the low...

Xin Wang, Zi-Jian Kang, Yu Cai et al. · 0 citations
Sep 2026

Multi-source geospatial fusion for enhanced visual geo-localization of unmanned aerial vehicles.

Currently, although significant advancements have been made in the autonomous navigation of Unmanned Aerial Vehicles (UAVs), reliable geo-localization remains a challenge when satellite navigation signals fail. Existing research primarily focuses on cross-view image retrieval, feature point matching, and visual localiz...

Xiong Qiu, Shou-Yi Liao, Dongfang Yang et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.