Text-to-hierarchical three-dimensional scene generation: a new approach for layered three-dimensional modeling from natural language
Rapid advancements in generative models have enabled substantial progress in text-to-3D scene synthesis. However, existing approaches often lack hierarchical structure and geometric consistency, limiting their use in complex, editable scene creation. To address this issue, an end-to-end hierarchical framework for text-to-3D scene generation is introduced. The main contribution of this study lies not in advancing a single generative model but in the design of a framework that synergistically integrates state-of-the-art components for video synthesis and mesh reconstruction. By decoupling scene semantics from dynamic camera trajectory control, the proposed system efficiently produces structured object-level editable three-dimensional scenes from natural language. Comprehensive experiments showed that the proposed integrated approach achieved superior scene quality, consistency, and editing flexibility compared with existing methods, offering a powerful solution for applications, such as virtual reality and digital twins.