Sep 2026· IEEE Transactions on Visualization and Computer Graphics· Vol PP, pp. 1-15· 0 citations
Medicine
TL;DR
RecPoint is introduced, a structure-aware framework for text-guided 3D shape generation that rectifies semantic drift by explicitly modeling local geometric relationships and proposes a test-time locality-aware semantic rectification mechanism that bridges modality gaps by grounding textual features in visually similar, structurally aligned image embeddings.
Abstract
Zero-shot text-to-3D generation has garnered growing interest through the integration of vision-language foundation models and 3D generative frameworks. However, existing methods often overlook the preservation of local geometric semantics critical for producing semantically aligned and structurally coherent 3D point clouds. Here, we introduce RecPoint, a structure-aware framework for text-guided 3D shape generation that rectifies semantic drift by explicitly modeling local geometric relationships. Central to our method is a geometry-aware graph transformer that encodes point-wise adjacency into graph signals, enabling fine-grained feature learning and robust shape reconstruction. Complementing this, we propose a test-time locality-aware semantic rectification mechanism that bridges modality gaps by grounding textual features in visually similar, structurally aligned image embeddings. Without relying on paired text-shape data, RecPoint achieves superior fidelity and semantic alignment across standard benchmarks, outperforming previous state-of-the-art approaches.
A comprehensive geometric-semantic fusion mechanism that resolves geometric noise and semantic ambiguity by explicitly utilizing semantic guidance and formulating 3D segmentation as solving point-and-set merging and partitioning problems, and an innovative manifold-distance-based point cloud refinement strategy.
This work proposes a segmentation framework guided by Contrastive Language–Image Pre-training (CLIP) that enriches sparse 3D tokens with vision–language semantic priors and introduces a decoupled CLIP-induced semantic residual that forms semantic-geometric attention biases for local window attention.
SpecBridge is introduced, a 3D-2D-Text pre-training framework that leverages CLIP priors as a foundational bridge to connect three modalities by synergizing spectral graph theory with transitive semantic learning.
Dong Wang, Jie Jiang, Wei-Dong Min et al.· Proceedings of the Thirty-Fi...· 0 citations
Bi-SGL is proposed, a Bidirectional, Spatially Grounded, and Language-informed framework for point cloud semantic scene completion that improves semantic completion accuracy among evaluated point cloud SSC baselines while using substantially fewer parameters than cascaded dense-fusion models.
Houda Saffi, N. Otberdout, A. El Fallah Seghrouchni· IEEE Access· 0 citations
This work proposes ViCo-SAM3, a Vision-Conditioned alignment framework designed for OVCOS, which introduces vision-conditioned (ViCo) module, which dynamically modulates text embeddings with global visual context, enabling textual representations to adapt to the current image content and thereby effectively bridges the...
Qiang-Qiang Zhou, Wengang Tang, Yong Chen et al.· 1 citation
Qwen-3D is introduced, a geometry-aware LMM that compresses visual information within the Qwen backbone using multi-view geometric cues, enabling efficient long-horizon visual reasoning over static scenes and incorporates a query-based segmentation decoder that grounds language directly in the underlying 3D scene repre...
Lucy Lin, Ayush Jain, Yifan Liu et al.· 3 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.