Skip to content
Preprint

Beyond Isolated Objects: Relationship-aware Open Vocabulary Scene Understanding via 3D Scene Graph Analysis

Jul 2026 · 0 citations · 54 references
Computer Science

TL;DR

An Adaptive Gated Dual-Stream Contextual GAT that separates dense geometric features and semantic CLIP embeddings, performs edge-guided message passing, and adaptively fuses complementary semantics is introduced that separates dense geometric features and semantic CLIP embeddings, and adaptively fuses complementary semantics.

Abstract

Open-vocabulary 3D scene understanding aims to segment 3D scenes beyond predefined categories by transferring semantic knowledge from vision-language models. Existing methods have advanced this task by lifting language-aligned 2D features into 3D, yet they often rely on context-independent semantic representations, leaving object relationships underexplored for contextual refinement. We propose RelGraphOV, a relationship-aware framework that uses 3D scene graphs to enhance open-vocabulary 3D understanding. Our method constructs relational scene graphs from multi-view observations by leveraging vision-language reasoning to infer object relationships and prune geometrically implausible connections, without manual relationship annotations. To aggregate relational context while avoiding feature interference, we introduce an Adaptive Gated Dual-Stream Contextual GAT that separates dense geometric features and semantic CLIP embeddings, performs edge-guided message passing, and adaptively fuses complementary semantics. A hierarchical contrastive objective further promotes instance-level consistency and category-level discrimination. Experiments on ScanNetV2, ScanNet200, ScanNet$++$, and Replica demonstrate strong performance and generalization ability. Project Page: https://cxavireh.github.io/relgraphov-projectpage

View source

Similar papers

Open access Jul 2026

KNA-SG: Keyframe–Node-Associated Open-Vocabulary 3D Scene Graphs from RGB Sequences

KNA-SG, a framework for constructing open-vocabulary 3D scene graphs from RGB sequences with explicit keyframe–node associations, is proposed and Experimental results show that KNA-SG outperforms existing methods on open-vocabulary 3D semantic segmentation and 3D object grounding tasks.

Yang Xu, Wen-ku Shi, Jing Xing et al. · 0 citations
Conference Jul 2026

3D vision-language question answering with explicit scene graphs and local topology priors

TA-LMM, a 3D visual question answering method built on explicit scene graphs and local topological priors, is proposed, suggesting that explicit local topological priors can improve scene consistency in 3D visual question answering.

Kaixin Wu, Kunlin Zhou, Boxin Li et al. · 0 citations
Open access 2026

Bi-SGL: Bidirectional, Spatially Grounded, and Language-Informed Framework for Semantic Scene Completion

Bi-SGL is proposed, a Bidirectional, Spatially Grounded, and Language-informed framework for point cloud semantic scene completion that improves semantic completion accuracy among evaluated point cloud SSC baselines while using substantially fewer parameters than cascaded dense-fusion models.

Houda Saffi, N. Otberdout, Amal El Fallah Seghrouchni · 0 citations
Preprint Aug 2026

GroupForward: Building Referable 3D Scenes via Instance-Grouped Feed-Forward Gaussian Splatting

This work proposes GroupForward, an instance-grouped feed-forward Gaussian splatting model that reconstructs geometry, appearance, instance structure, and semantics from sparse, unposed, and uncalibrated multi-view images and proposes a Referential Scene Reasoning Framework (RSRF) for complex 3D referring segmentation.

Qijian Tian, Zimeng Wu, Xuhong Wang et al. · 0 citations
Preprint Aug 2026

GuideGround: VLM-guided Semantic Understanding and Viewpoint-aware Reasoning for 3D Visual Grounding

3D visual grounding aims to localize the target object in a 3D scene from a natural language query, requiring both fine-grained semantic understanding and viewpoint-dependent spatial reasoning. Existing methods typically formulate semantic understanding as an auxiliary closed-set object classification task and rely on multi-view feature aggregation for viewpoint reasoning, limiting semantic generalization and weakening viewpoint-specific evidence. We observe that vision-language models naturally provide complementary capabilities through open-vocabulary semantic understanding and global scene perception. Based on this insight, we propose GuideGround, a VLM-guided framework that complements rather than replaces task-specific grounding models by leveraging VLMs for semantic enhancement and viewpoint-specific hypothesis verification. Specifically, we replace auxiliary closed-set object classification with VLM-generated object semantic descriptions to enhance semantic understanding. Meanwhile, instead of directly aggregating multi-view representations, we preserve viewpoint-specific grounding hypotheses through per-view grounding and explicitly verify them using VLMs across candidate viewpoints. Extensive experiments on the ReferIt3D benchmark demonstrate that GuideGround consistently outperforms previous state-of-the-art methods. Comprehensive ablation studies further confirm the effectiveness of both the proposed semantic understanding and viewpoint reasoning strategies.

Yiwen Wang, Yuyang Deng, Yihao Long et al. · 0 citations
Aug 2026

Hyperbolic Hierarchical 3D Scene Representations for Open-Vocabulary 3D Scene Graph Generation.

Open-vocabulary 3D scene graph generation aims to predict 3D objects and their predicates beyond the annotated label space. Compared to closed-set 3D scene graph generation methods, the open-vocabulary approach is more general, practical, and less dependent on labor-intensive ground truth annotations. Existing open-vocabulary 3D scene graph generation methods rely on learning individual object and predicate features in the representation space while ignoring higher-level 3D scene representations, leading to overfitting and suboptimal performance. In this work, we propose a hyperbolic learning-based approach to address this problem by leveraging hyperbolic geometry to learn hierarchical 3D scene representations in the form of scene-region-instance, where the scene represents the complete 3D environment, a region contains related instances, and an instance corresponds to an individual object or predicate. Specifically, our method decomposes a 3D scene into a discrete hierarchy consisting of scene, region, and instance nodes, and embeds this hierarchy into a learned hyperbolic representation space. The learned hyperbolic embeddings are optimized in a bottom-up manner, where higher-level nodes are derived from their corresponding child nodes. The learned hierarchical 3D scene representations are incorporated as structural and semantic guidance for open-vocabulary 3D scene graph generation. We further observe that outliers in the form of erroneous hyperbolic embeddings can negatively impact hierarchical reasoning. To mitigate their negative impact, we present an enhancement strategy that learns an adaptive distance metric robust to the outliers over the learned hyperbolic representation space and subsequently improves overall performance. Extensive experiments on 3DSSG and ScanNet datasets demonstrate the effectiveness of our method in 3D scene graph generation under closed-set, open-vocabulary, and zero-shot settings.

Haoran Hou, Mingtao Feng, Qing Zhu et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.