DenseRS-CLIP: Enhancing Dense Feature Representation of Remote Sensing CLIP via Attention-Decoupled Dual-Branch Distillation
Existing remote sensing vision–language foundation models mainly follow the CLIP-style global image-text alignment paradigm. While effective for image-level semantic understanding, this paradigm leaves patch-level representations insufficiently discriminative and spatially inconsistent for dense prediction tasks. To ad...