Stage-Adaptive Spatial-Frequency Decomposition and Enhancement With Multimodal Conditional Routing for Remote Sensing Image Segmentation
Abstract
Multimodal remote sensing image segmentation benefits from complementary cues provided by heterogeneous data sources, but accurate segmentation remains challenging due to difficulties in effectively fusing multimodal features and preserving fine-grained structures. Although spatial-frequency fusion methods have shown promise, existing approaches often rely on fixed or input-agnostic frequency partitioning and overlook the different objectives of encoding and decoding stages, limiting their adaptability to scene-dependent frequency distributions and stage-specific representation needs. To address these limitations, we propose a stage-adaptive spatial-frequency decomposition and enhancement network (SF-DENet) for multimodal remote sensing image segmentation. SF-DENet adopts a two-stream encoder–decoder architecture and introduces two stage-specific spatial-frequency fusion modules with distinct interaction patterns. During encoding, the encoder spatial-frequency fusion (EnFusion) module feeds the fused optical–auxiliary features into parallel spatial and frequency branches and employs adaptive frequency masks to modulate the amplitude spectrum, thereby enhancing cross-modal semantic alignment. During decoding, the decoder spatial-frequency fusion (DeFusion) module introduces explicit high- and low-frequency cues guided by the original inputs and performs frequency-aware interaction between decoder features and skip-connected encoder features, thereby enhancing structure-aware detail reconstruction. To overcome fixed or input-agnostic frequency partitioning, the frequency branch incorporates a conditional composition-based frequency decomposition (C$^{2}$FD) module, which predicts routing weights from multimodal inputs and composes multiple soft-edge frequency-mask experts into content-adaptive masks for frequency representation modulation. Experiments on multiple multimodal remote sensing benchmarks demonstrate that SF-DENet achieves superior segmentation performance, particularly in scenes with complex textures, ambiguous boundaries, and modality inconsistency.