An Adaptive Gated Dual-Stream Contextual GAT that separates dense geometric features and semantic CLIP embeddings, performs edge-guided message passing, and adaptively fuses complementary semantics is introduced that separates dense geometric features and semantic CLIP embeddings, and adaptively fuses complementary semantics.
Abstract
Open-vocabulary 3D scene understanding aims to segment 3D scenes beyond predefined categories by transferring semantic knowledge from vision-language models. Existing methods have advanced this task by lifting language-aligned 2D features into 3D, yet they often rely on context-independent semantic representations, leaving object relationships underexplored for contextual refinement. We propose RelGraphOV, a relationship-aware framework that uses 3D scene graphs to enhance open-vocabulary 3D understanding. Our method constructs relational scene graphs from multi-view observations by leveraging vision-language reasoning to infer object relationships and prune geometrically implausible connections, without manual relationship annotations. To aggregate relational context while avoiding feature interference, we introduce an Adaptive Gated Dual-Stream Contextual GAT that separates dense geometric features and semantic CLIP embeddings, performs edge-guided message passing, and adaptively fuses complementary semantics. A hierarchical contrastive objective further promotes instance-level consistency and category-level discrimination. Experiments on ScanNetV2, ScanNet200, ScanNet$++$, and Replica demonstrate strong performance and generalization ability. Project Page: https://cxavireh.github.io/relgraphov-projectpage
KNA-SG, a framework for constructing open-vocabulary 3D scene graphs from RGB sequences with explicit keyframe–node associations, is proposed and Experimental results show that KNA-SG outperforms existing methods on open-vocabulary 3D semantic segmentation and 3D object grounding tasks.
Yang Xu, Wen-ku Shi, Jing Xing et al.· Technologies· 0 citations
TA-LMM, a 3D visual question answering method built on explicit scene graphs and local topological priors, is proposed, suggesting that explicit local topological priors can improve scene consistency in 3D visual question answering.
Kaixin Wu, Kunlin Zhou, Boxin Li et al.· International Conference on...· 0 citations
Bi-SGL is proposed, a Bidirectional, Spatially Grounded, and Language-informed framework for point cloud semantic scene completion that improves semantic completion accuracy among evaluated point cloud SSC baselines while using substantially fewer parameters than cascaded dense-fusion models.
Houda Saffi, N. Otberdout, Amal El Fallah Seghrouchni· IEEE Access· 0 citations
This work proposes GroupForward, an instance-grouped feed-forward Gaussian splatting model that reconstructs geometry, appearance, instance structure, and semantics from sparse, unposed, and uncalibrated multi-view images and proposes a Referential Scene Reasoning Framework (RSRF) for complex 3D referring segmentation.
Qijian Tian, Zimeng Wu, Xuhong Wang et al.· 0 citations
3D visual grounding aims to localize the target object in a 3D scene from a natural language query, requiring both fine-grained semantic understanding and viewpoint-dependent spatial reasoning. Existing methods typically formulate semantic understanding as an auxiliary closed-set object classification task and rely on multi-view feature aggregation for viewpoint reasoning, limiting semantic generalization and weakening viewpoint-specific evidence. We observe that vision-language models naturally provide complementary capabilities through open-vocabulary semantic understanding and global scene perception. Based on this insight, we propose GuideGround, a VLM-guided framework that complements rather than replaces task-specific grounding models by leveraging VLMs for semantic enhancement and viewpoint-specific hypothesis verification. Specifically, we replace auxiliary closed-set object classification with VLM-generated object semantic descriptions to enhance semantic understanding. Meanwhile, instead of directly aggregating multi-view representations, we preserve viewpoint-specific grounding hypotheses through per-view grounding and explicitly verify them using VLMs across candidate viewpoints. Extensive experiments on the ReferIt3D benchmark demonstrate that GuideGround consistently outperforms previous state-of-the-art methods. Comprehensive ablation studies further confirm the effectiveness of both the proposed semantic understanding and viewpoint reasoning strategies.
Yiwen Wang, Yuyang Deng, Yihao Long et al.· 0 citations
Open-vocabulary 3D scene graph generation aims to predict 3D objects and their predicates beyond the annotated label space. Compared to closed-set 3D scene graph generation methods, the open-vocabulary approach is more general, practical, and less dependent on labor-intensive ground truth annotations. Existing open-vocabulary 3D scene graph generation methods rely on learning individual object and predicate features in the representation space while ignoring higher-level 3D scene representations, leading to overfitting and suboptimal performance. In this work, we propose a hyperbolic learning-based approach to address this problem by leveraging hyperbolic geometry to learn hierarchical 3D scene representations in the form of scene-region-instance, where the scene represents the complete 3D environment, a region contains related instances, and an instance corresponds to an individual object or predicate. Specifically, our method decomposes a 3D scene into a discrete hierarchy consisting of scene, region, and instance nodes, and embeds this hierarchy into a learned hyperbolic representation space. The learned hyperbolic embeddings are optimized in a bottom-up manner, where higher-level nodes are derived from their corresponding child nodes. The learned hierarchical 3D scene representations are incorporated as structural and semantic guidance for open-vocabulary 3D scene graph generation. We further observe that outliers in the form of erroneous hyperbolic embeddings can negatively impact hierarchical reasoning. To mitigate their negative impact, we present an enhancement strategy that learns an adaptive distance metric robust to the outliers over the learned hyperbolic representation space and subsequently improves overall performance. Extensive experiments on 3DSSG and ScanNet datasets demonstrate the effectiveness of our method in 3D scene graph generation under closed-set, open-vocabulary, and zero-shot settings.
Haoran Hou, Mingtao Feng, Qing Zhu et al.· IEEE Transactions on Pattern...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.