Sep 2026· Frontiers in Artificial Intelligence· 0 citations· 55 references
TL;DR
This work proposes a segmentation framework guided by Contrastive Language–Image Pre-training (CLIP) that enriches sparse 3D tokens with vision–language semantic priors and introduces a decoupled CLIP-induced semantic residual that forms semantic-geometric attention biases for local window attention.
Abstract
Recent 3D Transformers have become a dominant framework for point-cloud segmentation by modeling spatial context in sparse 3D scenes. However, geometry and color alone provide limited high-level semantic cues, especially for cluttered boundary regions, visually similar objects, and long-tail categories. To address this issue, we propose a segmentation framework guided by Contrastive Language–Image Pre-training (CLIP) that enriches sparse 3D tokens with vision–language semantic priors. Specifically, dense CLIP features extracted from multi-view RGB images are projected onto 3D points through visibility-aware alignment and view pooling, and are fused with relative geometric offsets and color cues to form semantically aware sparse voxel tokens. To better exploit the aligned CLIP semantics during local token interactions, we build on contextual relative signal encoding (cRSE) and introduce a decoupled CLIP-induced semantic residual that forms semantic-geometric attention biases for local window attention. We further adapt block-wise online softmax computation to generate and consume these biases on the fly. Experiments on ScanNet, ScanNet200, and S3DIS demonstrate competitive segmentation performance, improved instance-level discrimination, and a 25.7% reduction in peak
online 3D-stage
training memory compared with the materialized attention implementation when cached CLIP features are used.
GaussianDS, a depth-supervised semantic 3DGS framework that treats semantic lifting as a supervision-alignment problem and jointly optimizes RGB appearance, rendered depth, and compact semantics from scratch, is proposed.
Vision-language models excel at 2D image understanding but remain limited in 3D spatial reasoning. Progress is hindered by limitations in current benchmarks. First, 3D datasets often rely on point clouds that capture geometry but discard rich visual features like texture, text, and materials. Second, annotations treat...
Anubhav Khanal, Prabigya Acharya, Roshni Poudel et al.· 0 citations
Unified 3D vision-language systems must combine complementary geometry, scale, and appearance cues while supporting tasks from instance segmentation to language-guided reasoning. Existing methods often process point clouds, voxel grids, and multi-view images independently; directly combining these heterogeneous represe...
Xue-Qi Qiu, Xing-Yu Miao, Jing-Jing Deng et al.· 0 citations
Results suggest that decoupled matching improves robustness to appearance-driven confusion in indoor point-cloud segmentation and propose D2M-Net, a decoupled dual-matching network that separates backbone features into geometry-oriented and semantic-oriented subspaces before prototype comparison.
Han-Bin Fang, Cheng-Long Peng, Xue-Yong Xiang et al.· The Visual Computer· 0 citations
This paper proposes a semantic parsing method that leverages building-structure priors that uses a shifted-window hierarchical transformer encoder to extract multi-scale visual features and combines a gradient-direction-consistency line segment detection algorithm to construct a Manhattan 3D bounding box.
A 3D Local- Global Linear Attention Mechanism (LG-LAM) is devised that efficiently captures long-range contextual information with linear complexity, enabling a comprehensive understanding of the 3D scene without heavy computational burdens.
Jie Li, Jia-Heng Xu, Laiyan Ding et al.· International Conference on...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.