Optimal Transport-Driven Visual Semantic Consensus for Remote Sensing Image Matching
Remote sensing image matching is a fundamental task for numerous Earth observation applications, for example, change detection, multitemporal analysis, and cross-sensor fusion. However, seeking reliable matching remains highly challenging due to substantial variations in imaging conditions. Existing approaches often suffer from limited robustness when confronted with large geometric deformations. To address this, we propose a novel optimal transport (OT)-guided semantic consensus framework. Unlike traditional consensus methods that rely on the spatial topology of feature points, our approach leverages visual semantic cues to enhance the discriminability of putative matching. Specifically, by integrating visual–language large models, we design a semantic consensus learning paradigm, which can be used to evaluate the visual semantic consistency of correspondences. More importantly, grounded in OT theory, our semantic consensus learning framework can be trained in a self-supervised manner without requiring additional manual annotations. In this work, we also provide a model instantiation that integrates the proposed semantic consensus learning with mismatch rejection networks. Overall, this work provides a general paradigm for integrating self-supervised learning with local semantic consensus learning. Comprehensive evaluations on multiple datasets show that our method not only outperforms state-of-the-art baselines in matching accuracy but also generalizes well to diverse image deformations.