Skip to content
Open access

CAF-Net: A Unified Framework for Resolving Spatial–Frequency Representation Conflicts in Multimodal Remote Sensing Segmentation

Jul 2026 · Remote Sensing · Vol 18, pp. 2284 · 0 citations · 26 references

TL;DR

The results indicate that explicitly modeling spatial–frequency discrepancies can improve multimodal segmentation accuracy and representation consistency.

Abstract

Multimodal remote sensing segmentation commonly integrates optical imagery with Digital Surface Models (DSMs) to improve land-cover understanding. However, existing methods often overlook a critical issue, namely the inconsistency between spatial-domain structural details and frequency-domain semantics, which can lead to misaligned features and degraded performance in complex scenes. To address this problem, we formulate spatial–frequency representation conflict as a unified multimodal learning problem and propose the Conflict-Aware Fusion Network (CAF-Net). Specifically, Cross-Modal Structure Guidance (CMSG) extracts DSM-derived high-pass structural cues and conditionally modulates optical features to improve boundary consistency. The Adaptive Cross-Frequency Module (ACFM) separates DCT coefficients using a fixed radial mask, adaptively reweights low- and high-frequency components, and performs cross-modal alignment at the highest encoder stage. Uncertainty-Aware Fusion (UAF) predicts pixel-wise relative reliability scores and normalizes them across modalities to suppress low-confidence responses. This coordinated design links spatial refinement, frequency alignment, and reliability-guided fusion instead of treating them as independent feature-enhancement operations. Experiments on the ISPRS Vaihingen and Potsdam datasets yield mIoU scores of 84.35% and 86.86%, respectively. The results indicate that explicitly modeling spatial–frequency discrepancies can improve multimodal segmentation accuracy and representation consistency.

Read PDF

Similar papers

Open access Aug 2026

GCF-Net: Stage-Aligned Optical–Elevation Fusion for Aerial Remote Sensing Semantic Segmentation

Optical–elevation data fusion is widely used in aerial remote sensing semantic segmentation, as optical imagery provides rich spectral and textural information, while DSM or DEM data offer complementary elevation-related structural cues. However, effective fusion remains challenging because optical and elevation representations may exhibit cross-modal structural inconsistency, frequency–spatial response imbalance, and decoder-stage structural attenuation. To address these challenges, we propose GCF-Net, a stage-aligned optical–elevation fusion network that matches different cross-modal processing objectives to the evolving representation states of the encoder–decoder pipeline. A Structure-Guided Cross-Modal Correction Module first performs structure-conditioned correction of modality-specific features before fusion. A Frequency–Spatial Cross-Modal Fusion Module then constructs joint representations through bounded cross-modal magnitude conditioning, frequency-to-spatial reconstruction, and spatial recalibration. During decoding, a Geometry-Aware Cross-Scale Refinement Module reintroduces elevation-derived structural guidance into multiscale fused features. Experiments on ISPRS Vaihingen, ISPRS Potsdam, and MMHunan yield mIoU scores of 72.37%, 75.38%, and 52.23%, respectively, achieving the highest mIoU among the evaluated unimodal, multimodal, and SAM-based methods under the unified protocol. Ablation and replacement experiments verify the complementary roles of the three stage-specific components, while sensitivity, elevation perturbation, and complexity analyses indicate architectural flexibility, tolerance to moderate elevation degradation, and a balanced accuracy–efficiency trade-off.

Yifan Yu, Song Deng, Yang Yang et al. · 0 citations
2026

A Coarse-to-Fine Progressive Fusion Network for Multimodal Semantic Segmentation of Remote Sensing Images

Multimodal semantic segmentation of high-resolution remote sensing imagery is important for fine-grained land-cover interpretation. However, existing fusion methods still suffer from unstable shallow optical-DSM alignment and deep feature degradation caused by heterogeneous frequency noise, boundary-detail loss, and inconsistent spatial responses. To address the aforementioned challenges, this letter proposes a coarse-to-fine progressive fusion network (CFPFNet). Specifically, a visual state space model extracts a Mamba-derived global structural prior to guide the coarse-grained context enhancement (CGCE) module for preliminary cross-modal alignment. Then, the fine-grained adaptive frequency-spatial fusion (FGAF) module performs amplitude-phase collaboration and adaptive spatial cross-gating for multiscale semantic refinement. Experiments on the ISPRS Vaihingen and Potsdam datasets demonstrate the effectiveness of CFPFNet. On Vaihingen, CFPFNet improves mIoU and mF1 by 2.22% and 1.39% over the simple dual-stream baseline, respectively.

Di Zhang, Yuhang Yan, Q. Niu et al. · 2 citations
Open access 2026

Stage-Adaptive Spatial-Frequency Decomposition and Enhancement With Multimodal Conditional Routing for Remote Sensing Image Segmentation

Multimodal remote sensing image segmentation benefits from complementary cues provided by heterogeneous data sources, but accurate segmentation remains challenging due to difficulties in effectively fusing multimodal features and preserving fine-grained structures. Although spatial-frequency fusion methods have shown promise, existing approaches often rely on fixed or input-agnostic frequency partitioning and overlook the different objectives of encoding and decoding stages, limiting their adaptability to scene-dependent frequency distributions and stage-specific representation needs. To address these limitations, we propose a stage-adaptive spatial-frequency decomposition and enhancement network (SF-DENet) for multimodal remote sensing image segmentation. SF-DENet adopts a two-stream encoder–decoder architecture and introduces two stage-specific spatial-frequency fusion modules with distinct interaction patterns. During encoding, the encoder spatial-frequency fusion (EnFusion) module feeds the fused optical–auxiliary features into parallel spatial and frequency branches and employs adaptive frequency masks to modulate the amplitude spectrum, thereby enhancing cross-modal semantic alignment. During decoding, the decoder spatial-frequency fusion (DeFusion) module introduces explicit high- and low-frequency cues guided by the original inputs and performs frequency-aware interaction between decoder features and skip-connected encoder features, thereby enhancing structure-aware detail reconstruction. To overcome fixed or input-agnostic frequency partitioning, the frequency branch incorporates a conditional composition-based frequency decomposition (C$^{2}$FD) module, which predicts routing weights from multimodal inputs and composes multiple soft-edge frequency-mask experts into content-adaptive masks for frequency representation modulation. Experiments on multiple multimodal remote sensing benchmarks demonstrate that SF-DENet achieves superior segmentation performance, particularly in scenes with complex textures, ambiguous boundaries, and modality inconsistency.

Xuran Pan, Rui Zhang, Zibo Xu et al. · 0 citations
2026

SCF-Net: A Flow-Guided Alignment-Enhanced SAM–CNN Hybrid Framework for Remote Sensing Change Detection

While integrating convolutional neural networks (CNNs) and the segment anything model (SAM) is promising for remote sensing change detection (CD), effectively synergizing them remains challenging. Existing hybrid methods often rely on simple feature concatenation, failing to bridge the gap between CNNs’ fine-grained local structures and SAM’s global semantic priors. Moreover, neglecting spatial misalignment in bi-temporal imagery often leads to pseudo-changes and boundary inconsistencies. To address these limitations, we propose SCF-Net, a flow-guided alignment-enhanced framework. First, a dual-path fusion module (DPFM) bridges the cross-modal semantic gap by embedding global contexts into local representations. Second, an optical-flow-guided differential enhancement module (OFDEM) implements adaptive flow-based warping to rectify spatial shifts. Finally, a cross-scale fusion module (CSFM) ensures hierarchical feature consistency. Extensive experiments across LEVIR-CD, CLCD, and GFSW-CLCD demonstrate that SCF-Net effectively mitigates granular mismatch and registration noise. The proposed model achieves competitive $F1$ -scores of 91.58%, 81.23%, and 83.80% on the three datasets, respectively, demonstrating its effectiveness and robustness compared to current mainstream algorithms.

Fan Yang, Yurong Qian, Xin Yang et al. · 0 citations
Open access Aug 2026

Frequency and Edge-Guided Segment Anything Model for Remote Sensing Image Semantic Segmentation

Frequency and Edge-guided SAM (FE-SAM) is proposed, a scalable and efficient framework for RSISS that adaptively decomposes and modulates frequency-domain features based on the input data and designs EGRefiner, which integrates multi-scale edge-enhanced information extracted from the input image.

Feng Gao, Zizhe Pan, Haoting Wang et al. · 0 citations
2026

Center-Aware Global-to-Local Modeling for Patch-Based HSI–LiDAR Classification

Hyperspectral imagery (HSI) and light detection and ranging (LiDAR) provide complementary spectral and elevation cues for fine-grained land-cover classification, yet accurate pixel-wise fusion remains challenging in heterogeneous scenes. Most deep HSI–LiDAR classifiers follow a center-supervised patch-based setting, where supervision is defined on the center pixel while predictions are inferred from a context-enhanced patch representation. Under this setting, mixed semantics within a patch induce heterogeneous-patch ambiguity: contextual pixels exert uneven influence on the center-pixel decision, and indiscriminate context aggregation can dilute center-consistent evidence in mixed and boundary regions. To address this issue, we propose a center-aware global-to-local refinement framework that explicitly regulates how contextual information is organized and accumulated under center supervision. First, a lightweight cross-modal channel alignment module fuses HSI and LiDAR features into a unified representation while preserving the spatial layout. Second, we introduce a center-aware global regulation module built on Vision Mamba, equipped with a dual-direction Spiral Scan that orders tokens in a periphery-to-center manner. This design induces a structured information flow that progressively consolidates center-relevant semantics under heterogeneous neighborhoods. Finally, a lightweight spatial–spectral refinement module (SSRM) refines discriminative local details within the globally regulated feature space by recovering boundary-sensitive structures and recalibrating channel responses. Extensive experiments on three public benchmarks (Houston2013, MUUFL, and Augsburg) demonstrate that the proposed method consistently outperforms representative local neighborhood modeling, within-patch global interaction, and local–global hybrid approaches. The code is available at https://github.com/lmwdhr/ViT–CNN

Mingwan Li, Sheng Fang, Zhe Li et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.