Skip to content

Author

Xueli Chang

1 paper indexed here

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

Sep 2026

StructPointNet: explicit geometric prior extraction and reuse for efficient multimodal remote sensing semantic segmentation

Semantic segmentation is widely regarded as one of the most advanced scene understanding techniques in remote sensing. Although monomodal segmentation models have advanced recently, the rich potential of multimodal data is still largely untapped. Current multimodal approaches are often rigid, typically limited to bimodal inputs, and failing to adapt flexibly when three or more co-registered two-dimensional grid-based modalities are available. A more pressing issue is the computational burden inherent in traditional fusion paradigms; as the number of input modalities increases, parameter counts and processing loads often increase substantially, making them less practical for scalable multi-source remote sensing interpretation. Thus, creating a scalable and lightweight framework for co-registered grid-based multimodal remote sensing data remains an open challenge. To bridge this gap, we introduce StructPointNet, a lightweight segmentation framework built on the philosophy of “explicit extraction and reuse of geometric priors.” Our approach balances efficiency with precision through three core mechanisms: the structure-sensitive modality encoder, which captures modality-specific high-frequency geometric details via parallel Sobel branches and edge-guided attention; the Heterogeneity rectification layer and pyramidal hybrid backbone, which map diverse features into a shared latent space for global-local context modeling; and the boundary-aware point decoder, which refines boundary segmentation by resampling shallow structural features based on uncertainty estimates. Thanks to these designs, StructPointNet functions as a flexible framework for monomodal and multimodal segmentation with variable numbers of co-registered grid-based inputs while keeping the additional cost of each lightweight modality-specific branch controlled compared with conventional multistream backbone replication. We benchmarked our approach against multiple representative state-of-the-art models from the last five years on two public multimodal datasets. Experimental results demonstrate that StructPointNet delivers competitive segmentation accuracy, particularly in defining complex geo-object boundaries. Moreover, the model proves to be scalable and hardware-friendly, offering a practical paradigm for efficient multimodal interpretation in resource-limited settings.

Siwei Wei, Ruoxi Wang, Xueli Chang · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.