To address the challenges of insufficient global semantic modeling and blurred boundaries in urban traffic scene segmentation, this study proposes a frequency-spatial collaborative framework based on DeepLabV3+. A Spectral Decoupling Adaptive Modulation (SDAM) module enhances low-frequency semantics and high-frequency details in the frequency domain. A Hierarchical Spatial Dependency Modeling (HSDM) module captures local consistency and global semantic dependencies, while a Structure-guided Adaptive Multi-scale Fusion (SAMF) module dynamically integrates multi-scale features using structural priors. Experiments on Cityscapes and CamVid demonstrate improvements of 3.6% and 1.6% over DeepLabV3+, respectively, while maintaining real-time performance.
Weiwei Zhao, Yingchao Dong, Lingchao Wang et al.· International Journal of Adv...· 0 citations
ABSTRACT Currently, existing point cloud semantic segmentation methods do not fully exploit surface geometric features. In particular, the depiction of object boundaries and the transition areas of curved surfaces is rather rough. On the other hand, the neighbourhood aggregation mostly follows a single strategy, making it difficult to simultaneously take into account the context and fine-grained differences and ignoring local details. To address these issues, this paper proposes a geometry-enhanced adaptive local feature aggregation network (GALA-Net). First, a geometric information embedding (GIE) module is introduced, which extracts pseudo-normal vectors and pseudo-curvatures of local point cloud regions as geometric priors, and incorporates multi-frequency sine–cosine encoding to capture multi-scale spatial relationships, yielding enhanced local geometric representations. Then, an adaptive feature fusion (AFF) module dynamically allocates fusion weights between semantic and geometric features, thereby alleviating channel coupling and neighbourhood noise amplification caused by simple concatenation. Next, a dual-path adaptive attention aggregation (DAAA) module jointly models semantic and positional attention and adaptively fuses them with max-pooled features to improve the robustness of local aggregation. In addition, a self-enhanced attention encoding (SEAE) module is designed to expand the feature representation space by extracting features through independent mapping branches and fusing them in a residual manner. The proposed model is evaluated on the S3DIS and ScanNetV2 datasets, achieving mIoU scores of 78.0% and 71.6%, respectively, which demonstrates its strong segmentation performance on indoor scenes.
Guiru Liu, Yingchao Dong, Lulin Wang et al.· International Journal of Rem...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.