Skip to content

Author

Mingtao Feng

1 paper indexed here

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

Aug 2026

Hyperbolic Hierarchical 3D Scene Representations for Open-Vocabulary 3D Scene Graph Generation.

Open-vocabulary 3D scene graph generation aims to predict 3D objects and their predicates beyond the annotated label space. Compared to closed-set 3D scene graph generation methods, the open-vocabulary approach is more general, practical, and less dependent on labor-intensive ground truth annotations. Existing open-vocabulary 3D scene graph generation methods rely on learning individual object and predicate features in the representation space while ignoring higher-level 3D scene representations, leading to overfitting and suboptimal performance. In this work, we propose a hyperbolic learning-based approach to address this problem by leveraging hyperbolic geometry to learn hierarchical 3D scene representations in the form of scene-region-instance, where the scene represents the complete 3D environment, a region contains related instances, and an instance corresponds to an individual object or predicate. Specifically, our method decomposes a 3D scene into a discrete hierarchy consisting of scene, region, and instance nodes, and embeds this hierarchy into a learned hyperbolic representation space. The learned hyperbolic embeddings are optimized in a bottom-up manner, where higher-level nodes are derived from their corresponding child nodes. The learned hierarchical 3D scene representations are incorporated as structural and semantic guidance for open-vocabulary 3D scene graph generation. We further observe that outliers in the form of erroneous hyperbolic embeddings can negatively impact hierarchical reasoning. To mitigate their negative impact, we present an enhancement strategy that learns an adaptive distance metric robust to the outliers over the learned hyperbolic representation space and subsequently improves overall performance. Extensive experiments on 3DSSG and ScanNet datasets demonstrate the effectiveness of our method in 3D scene graph generation under closed-set, open-vocabulary, and zero-shot settings.

Haoran Hou, Mingtao Feng, Qing Zhu et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.