DenseRS-CLIP: Enhancing Dense Feature Representation of Remote Sensing CLIP via Attention-Decoupled Dual-Branch Distillation
Abstract
Existing remote sensing vision–language foundation models mainly follow the CLIP-style global image-text alignment paradigm. While effective for image-level semantic understanding, this paradigm leaves patch-level representations insufficiently discriminative and spatially inconsistent for dense prediction tasks. To address this issue, we propose DenseRS-CLIP, a region-structured two-stage adaptation framework for enhancing dense representations of remote sensing CLIP models. In the first stage, we perform dual-granularity image-text contrastive pretraining to obtain a domain-adapted CLIP model with robust global semantic alignment. In the second stage, we introduce an attention-decoupled dual-branch distillation framework that reuses existing bounding-box and mask-derived localization annotations to construct ROI-level distillation signals without requiring additional region-text descriptions. Specifically, dense features are decoupled into content and context branches. The content branch is optimized by region-structured semantic distillation with a region correlation constraint, which improves local semantic discriminability and suppresses interregion feature homogenization. The context branch is optimized by DINOv3-guided topology distillation, which aligns patch-level self-similarity structures to improve spatial consistency and boundary awareness. Experiments on seven benchmarks covering region classification, visual grounding, and referring image segmentation show that DenseRS-CLIP achieves consistent improvements over representative remote sensing VLFMs and controlled distillation variants. These results validate the effectiveness of weakly localized, branch-specific distillation for dense remote sensing vision–language representation learning.