Skip to content

A Coarse-to-Fine Progressive Fusion Network for Multimodal Semantic Segmentation of Remote Sensing Images

2026 · IEEE Geoscience and Remote Sensing Letters · Vol 23, pp. 7506705-7506705 · 2 citations · 14 references

Abstract

Multimodal semantic segmentation of high-resolution remote sensing imagery is important for fine-grained land-cover interpretation. However, existing fusion methods still suffer from unstable shallow optical-DSM alignment and deep feature degradation caused by heterogeneous frequency noise, boundary-detail loss, and inconsistent spatial responses. To address the aforementioned challenges, this letter proposes a coarse-to-fine progressive fusion network (CFPFNet). Specifically, a visual state space model extracts a Mamba-derived global structural prior to guide the coarse-grained context enhancement (CGCE) module for preliminary cross-modal alignment. Then, the fine-grained adaptive frequency-spatial fusion (FGAF) module performs amplitude-phase collaboration and adaptive spatial cross-gating for multiscale semantic refinement. Experiments on the ISPRS Vaihingen and Potsdam datasets demonstrate the effectiveness of CFPFNet. On Vaihingen, CFPFNet improves mIoU and mF1 by 2.22% and 1.39% over the simple dual-stream baseline, respectively.

View source

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.