GRE-Diff, a controllable and interactive diffusion-based framework that automates the creation and editing of apartment floor plans under user-specified constraints, is proposed, offering a practical step toward bridging AI-driven automation and human creativity in spatial design.
Abstract
Designing functional and aesthetically coherent floor plans requires exploring a vast space of possible room arrangements, a task that quickly becomes overwhelming for human designers. In this paper, we propose GRE-Diff, a controllable and interactive diffusion-based framework that automates the creation and editing of apartment floor plans under user-specified constraints. By combining AI-generated suggestions with real-time, human-in-the-loop editing, the system enables users to specify room types, room counts, boundary shapes, and editing operations through LLM-parsed instructions or GUI-based interaction. It then generates a diverse set of plausible and well-structured designs for refinement. At the core of our approach is Gaussian Room Embedding (GRE), a continuous latent representation that models each room as a spatial Gaussian distribution capturing its location and extent. Extensive experiments on the RPLAN dataset show that GRE-Diff produces high-quality, constraint-aware, and editable polygonal layouts, offering a practical step toward bridging AI-driven automation and human creativity in spatial design.
The designs produced by Mise-en-Sc\`ene are the closest to the ground truth in perceived quality among all compared methods, by a wide margin over both an LLM layout planner and a specialized layout transformer, while the match-and-place stage bridges the remaining fidelity gap to the ground-truth composites.
Zipeng Xu, Ryan Murdock, Umberto Michieli· 0 citations
Understanding and generating spatially coherent layouts from natural language remains a fundamental yet challenging task for large language models (LLMs). Existing LLMs often struggle to capture explicit geometric relationships and structural dependencies between objects. To address this issue, we propose SG-Layout, a graph-guided layout generation framework that explicitly incorporates structured spatial knowledge into LLMs. SG-Layout follows a two-stage training paradigm: (1) a graph-language feature alignment stage, where a relational graph encoder and a projector are trained to map scene-graph embeddings into the LLM's linguistic space; and (2) an instruction tuning stage, where LoRA-based adapters enable efficient fine-tuning for instruction-driven layout generation while keeping the backbone frozen. We evaluate SG-Layout on image layout generation, indoor scene synthesis and robotic object rearrangement tasks. Experimental results show that SG-Layout improves spatial reasoning accuracy and geometric consistency over the compact open-source backbone, with particularly clear advantages in relation-dense and compositionally complex scenes. These results highlight the effectiveness of graph-structured feature alignment for enhancing controllable layout generation.
Junsheng Wang, Chao Chen, Mengying Xie et al.· 0 citations
A post-training paradigm that integrates multimodal alignment (MA) and structural perception (SP) is proposed, MA enhances element interpretation by grounding metadata-defined elements to their visual counterparts, mitigating semantic drift, and SP models layer-aware inter-element spatial relationships to improve hierarchical understanding and reduce structural ambiguity.
Yiyang Huang, Zhaowen Wang, Simon Jenni et al.· 0 citations
Comprehensive evaluations on 12 large-scale graphs demonstrate that the proposed negative sampling-based algorithm outperforms state-of-the-art algorithms in neighborhood preservation and cluster separation and reduces memory consumption by 72% on average.
Xin Chen, Shuowei Hou, Yifan Wang et al.· 0 citations
This work presents StructuredEdit, a pipeline that reframes design editing as parameter manipulation rather than pixel generation and embeds hard design constraints into vision-language model fine-tuning by backpropagating pixel-level constraint violations through a lightweight differentiable rasterizer.
This work introduces SON-1K, a comprehensive benchmark for text-to-image generation, and proposes a new approach, the enhanced LMDpp, enhancing the performance of the novel two-stage Large Language Model (LLM)-grounded diffusion model pipeline (LMD).
Weiyue Li, Yi Li, Xiaoyue Wang et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.