PromptRefine: source-free unsupervised domain adaptation for remote sensing image classification via CLIP filtering
Abstract
Despite the success of deep learning in remote sensing (RS) image classification, substantial domain shifts—stemming from heterogeneous sensors and diverse environmental conditions—frequently compromise model reliability. Although source-free unsupervised domain adaptation (SFUDA) has emerged as a critical paradigm to bypass data privacy and storage constraints, existing methods remain fragile in complex RS scenes where noisy pseudo-labels often trigger catastrophic semantic drift. We propose PromptRefine, a white-box SFUDA framework designed to anchor target adaptation through cross-modal intelligence. Specifically, we leverage the zero-shot semantic priors of large-scale vision–language models (e.g., contrastive language–image pre-training) to rectify source-biased predictions via a dynamic prompt fine-tuning mechanism. The framework executes a three-stage alternating optimization strategy that integrates cross-modal semantic alignment, hard-sample mining via sliced Wasserstein discrepancy, and fine-grained prompt evolution. Evaluated across 18 cross-domain tasks on five benchmarks (UCM, WHU-RS19, AID, RSSCN7, and NWPU-RESISC45), PromptRefine consistently achieves remarkable performance. Specifically, it outperforms the leading vision transformer-based baseline (VisTA) by 0.42% and 0.47% on the two groups of cross-domain tasks and surpasses the top ResNet-based baseline (SRKT/DFENet) by 2.94% and 5.02% in average classification accuracy. The proposed method also outperforms other SFUDA and unsupervised domain adaptation methods on all 18 tasks, demonstrating its superior adaptation capability. Our approach provides a robust, privacy-preserving, and computationally efficient solution, setting a benchmark for scalable RS scene characterization.