DEViCo: Discriminative Evidence View Correspondence Learning for Drone-Satellite Cross-View Geo-Localization
Abstract
The core objective of cross-view geo-localization (CVGL) is to determine the location of a query image by retrieving its geographically corresponding reference image from a geo-tagged database. A fundamental challenge stems from substantial appearance and geometric discrepancies induced by the distinct imaging geometries of near-nadir satellite imagery and lower altitude oblique drone imagery. Under uniform full image feature extraction, discriminative evidence can be highly intermixed with view-specific contextual regions in the learned representations. To address this, we propose DEViCo: discriminative evidence view (DEV) correspondence learning for drone–satellite CVGL, which constructs a drone-side DEV and introduces a triview dual correspondence learning framework for drone–satellite matching. In particular, a Cross-View Semantic Consistency Attention Generation Module (CAGM) exploits cross-view semantic consistency to generate attention heatmaps that highlight discriminative evidence in drone images. A Discriminative Evidence View Construction Module (DEVCM) then adaptively processes these heatmaps to construct a compact and coherent DEV. Based on the original satellite view, original drone view, and drone DEV, DEViCo performs triview dual correspondence learning, with one correspondence preserving global context and the other emphasizing discriminative evidence. Importantly, DEViCo can be integrated with most visual backbones without modifying their internal architectures. Extensive experiments on University-1652 and SUES-200 demonstrate consistent improvements across multiple backbones, and DEViCo further boosts MEAN to achieve 96.72 R@1 on University-1652 and competitive state-of-the-art-level performance on both benchmarks.