Skip to content
Open access

Vision–language guided semantic-geometric transformer for memory-efficient 3D scene understanding

Sep 2026 · Frontiers in Artificial Intelligence · 0 citations · 55 references

TL;DR

This work proposes a segmentation framework guided by Contrastive Language–Image Pre-training (CLIP) that enriches sparse 3D tokens with vision–language semantic priors and introduces a decoupled CLIP-induced semantic residual that forms semantic-geometric attention biases for local window attention.

Abstract

Recent 3D Transformers have become a dominant framework for point-cloud segmentation by modeling spatial context in sparse 3D scenes. However, geometry and color alone provide limited high-level semantic cues, especially for cluttered boundary regions, visually similar objects, and long-tail categories. To address this issue, we propose a segmentation framework guided by Contrastive Language–Image Pre-training (CLIP) that enriches sparse 3D tokens with vision–language semantic priors. Specifically, dense CLIP features extracted from multi-view RGB images are projected onto 3D points through visibility-aware alignment and view pooling, and are fused with relative geometric offsets and color cues to form semantically aware sparse voxel tokens. To better exploit the aligned CLIP semantics during local token interactions, we build on contextual relative signal encoding (cRSE) and introduce a decoupled CLIP-induced semantic residual that forms semantic-geometric attention biases for local window attention. We further adapt block-wise online softmax computation to generate and consume these biases on the fly. Experiments on ScanNet, ScanNet200, and S3DIS demonstrate competitive segmentation performance, improved instance-level discrimination, and a 25.7% reduction in peak online 3D-stage training memory compared with the materialized attention implementation when cached CLIP features are used.

Read PDF

Similar papers

Preprint Aug 2026

GaussianDS: Depth-supervised Semantic Gaussian Splatting for Scene Understanding

GaussianDS, a depth-supervised semantic 3DGS framework that treats semantic lifting as a supervision-alignment problem and jointly optimizes RGB appearance, rendered depth, and compact semantics from scratch, is proposed.

Yu-Fei Zhang, Chen-Lu Zhan, Hong-Wei Wang · 0 citations
#artificial intelligence Preprint Sep 2026

SceneBench: A Hierarchical Benchmark for Vision-Language Understanding of 3D Scenes

Vision-language models excel at 2D image understanding but remain limited in 3D spatial reasoning. Progress is hindered by limitations in current benchmarks. First, 3D datasets often rely on point clouds that capture geometry but discard rich visual features like texture, text, and materials. Second, annotations treat...

Anubhav Khanal, Prabigya Acharya, Roshni Poudel et al. · 0 citations
Preprint Sep 2026

From Alignment to Fusion in 3D Vision-Language

Unified 3D vision-language systems must combine complementary geometry, scale, and appearance cues while supporting tasks from instance segmentation to language-guided reasoning. Existing methods often process point clouds, voxel grids, and multi-view images independently; directly combining these heterogeneous represe...

Xue-Qi Qiu, Xing-Yu Miao, Jing-Jing Deng et al. · 0 citations
Aug 2026

D2M-Net: decoupled dual matching for few-shot 3D point-cloud semantic segmentation

Results suggest that decoupled matching improves robustness to appearance-driven confusion in indoor point-cloud segmentation and propose D2M-Net, a decoupled dual-matching network that separates backbone features into geometry-oriented and semantic-oriented subspaces before prototype comparison.

Han-Bin Fang, Cheng-Long Peng, Xue-Yong Xiang et al. · 0 citations
Aug 2026

A Semantic Parsing Method for Indoor Scene Images Based on Prior Knowledge of Building Structure.

This paper proposes a semantic parsing method that leverages building-structure priors that uses a shifted-window hierarchical transformer encoder to extract multi-scale visual features and combines a gradient-direction-consistency line segment detection algorithm to construct a Manhattan 3D bounding box.

Honglin Zhou, Songyang Ding, Jin-Tao Jiang · 0 citations
Conference Aug 2026

Enhancing 3D semantic scene completion via efficient attention and feature augmentation

A 3D Local- Global Linear Attention Mechanism (LG-LAM) is devised that efficiently captures long-range contextual information with linear complexity, enabling a comprehensive understanding of the 3D scene without heavy computational burdens.

Jie Li, Jia-Heng Xu, Laiyan Ding et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.