Skip to content
Open access

Text-Guided Semantic Segmentation Method for Indoor 3D Point Clouds

Aug 2026 · The International Archives of the Photogrammetry, Remote Sensing and Spatial Information Sciences · 0 citations · 14 references

TL;DR

Results indicate that incorporating textual semantic priors can effectively enhance high-level semantic representations of point clouds, providing a feasible solution for indoor 3D scene understand.

Abstract

Abstract. Point cloud semantic segmentation of indoor environments is a fundamental task in 3D scene understanding. However, existing methods mainly rely on geometric structures and color information, which are prone to error results in scenarios involving occlusion, sparse sampling, and geometrically similar structures. To address this issue, this paper proposes a text-knowledge-guided method for the point cloud semantic segmentation of indoor 3D scene. Built upon RandLA-Net as the baseline, the proposed method first constructs the textual semantic prototypes using multi-template prompts, and further enhances the stability of semantic anchors through periodic prototype refreshing. Then, a cross-modal semantic feature alignment mechanism is introduced at both the shallow and the high-level feature stages. Through feature alignment, bidirectional semantic interaction, and gated fusion, textual priors are progressively injected into the point cloud feature learning process. Finally, the model is jointly trained with a point-wise classification loss, a text-prototype alignment constraint, and a boundary optimization constraint to improve the semantic feature discrimination and segmentation boundary quality. Experimental results on the S3DIS dataset demonstrate that the proposed method achieves 86.8% OA, 81.6% mAcc, and 67.2% mIoU, exhibiting more stable segmentation performance in complex indoor scenes. As a consequence, these results indicate that incorporating textual semantic priors can effectively enhance high-level semantic representations of point clouds, providing a feasible solution for indoor 3D scene understand.

Read PDF

Similar papers

Open access Aug 2026

PCMTRefer: Symmetry-Aware Text-Guided Point Mamba for Referring Segmentation in Indoor 3D Point Clouds

Indoor 3D referring segmentation aims to identify and segment the target object in a point cloud according to a natural language expression. Although recent advances in multimodal feature fusion have markedly improved this task, existing methods do not jointly model the structural regularities commonly observed in indoor objects, such as approximate symmetry and repetitive local patterns, in conjunction with the directional constraints conveyed by referring expressions. In this work, we formulate symmetry-aware representation as the extraction of direction-consistent structural responses from both forward and backward traversals of the same spatially ordered sequence while preserving direction-sensitive variations introduced by occlusion, point cloud incompleteness, and cluttered scene layouts. Based on this formulation, we propose PCMTRefer, a symmetry-aware text-guided Point Mamba framework for indoor 3D referring segmentation. The input point cloud is first partitioned using an octree and arranged into a spatially coherent sequence via Z-order (Morton) ordering. A bidirectional state-space encoder then aggregates complementary context from both traversal directions, yielding richer representations of regular boundaries, repetitive structures, and approximately bilateral object geometries. In parallel, an asymmetric text-to-point guidance module injects semantic cues—object categories, attributes, and spatial relationships—into point-wise features, while a Background-Relaxation Token offers an auxiliary matching channel for non-target background regions. A Gumbel-Softmax-based semantic primitive learning module further extracts discriminative cues from referring expressions and integrates language semantics with point-level geometric features through a multi-scale decoder. Experimental results on the ScanRefer benchmark show that PCMTRefer achieves an Overall Acc@0.25 of 58.56%, an Overall Acc@0.5 of 54.19%, and an mIoU of 49.97%. Multi-seed validation further confirms the statistical stability of these results, with standard deviations of less than 0.2% across three independent runs.

Li Yuan, Bo Kong, Chenhao Li et al. · 0 citations
Open access Jul 2026

Unifying Street Scene Point Cloud Semantic Segmentation with Deformable Mesh-based Neural Representation

Abstract. Accurate semantic segmentation of urban point clouds is important for applications such as urban planning and autonomous driving. Recently, neural scene representations have been extended to merge semantic information across modalities and spatial dimensions. While 3D Gaussian Splatting (3DGS) enables efficient and high-quality reconstruction, its semantic understanding performance in street scenes is influenced by trajectory-constrained viewpoints, where Gaussian densification introduces occlusions and semantic ambiguity. This paper explores the use of NeRF-based neural representation for street scene point cloud semantic segmentation. Specifically, deformable neural mesh primitives (DNMPs) are used to compactly represent spatial geometry and simplify ray sampling. Then, neural fields including density, RGB, and semantics are constructed based on mesh vertex feature interpolation and MLPs. The sampled neural field values are accumulated via ray rendering and supervised using original images and corresponding semantic label maps generated by pre-trained models. Point cloud semantics are then predicted by interpolating neighboring samples within the learned field. The method is validated on the KITTI-360 and Waymo datasets. Results show that the proposed approach achieves improved semantic segmentation performance while maintaining competitive rendering quality, and supports both novel view synthesis and semantic rendering.

Yuzhou Zhou · 0 citations
Open access Aug 2026

Boundary Cues for Improved 3D Semantic Segmentation

A lightweight boundary-aware learning framework that explicitly models boundary regions during training is proposed, showing that incorporating boundary-aware supervision provides an effective and efficient approach to improving segmentation quality in challenging regions.

Waseem Iqbal, J. Paffenholz · 0 citations
Jul 2026

Mask-aware tri-modal learning for indoor 3D object detection

A mask-aware tri-modal framework that improves the quality of superpoint representations by retrieving a scene-level structural context from a pretrained PointSAM encoder to enhance object-centric evidence and predicting a soft mask weight to suppress unreliable superpoints.

Feng Zhou, Hui Wang, Kaida Ning et al. · 0 citations
Open access Aug 2026

Geometry-Enhanced Point–Voxel Fusion with Active Learning for Label-Efficient Point Cloud Semantic Segmentation

Point cloud semantic segmentation is fundamental for 3D scene understanding and has been widely used in autonomous driving and infrastructure inspection applications. However, its performance is often limited by insufficient representation of local geometric structures and the high cost of point-wise annotation. To address these issues, this paper proposes GeoFuse-AL, a label-efficient segmentation framework that integrates a geometry-guided point–voxel network with multi-cue active learning. The base model, GeoFuseNet, builds on a hybrid point–voxel backbone and incorporates a Local Geometry Prototype Attention module to enhance object boundaries, fine-grained structures, and local geometric patterns. An Adaptive Channel Fusion module is further designed to improve feature interaction between point-level details and voxel-level context. To reduce annotation dependence, a Multi-Cue Diversity Active Sampling strategy combines prediction uncertainty, color-gradient variation, geometric curvature, and feature-space clustering to select informative and diverse samples. Experiments on S3DIS and SemanticKITTI demonstrate that the proposed model achieves mIoU scores of 63.3% and 61.7%, respectively, outperforming several representative methods. Under limited annotation settings, the proposed strategy reaches 99.2% of fully supervised performance with only 15% labeled data on S3DIS and 97.6% with only 5% labeled data on SemanticKITTI. These results demonstrate that GeoFuse-AL improves segmentation accuracy while substantially reducing annotation requirements.

Cheng Zhang, Fei Meng, Yi-Chang Qiu et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.