The designs produced by Mise-en-Sc\`ene are the closest to the ground truth in perceived quality among all compared methods, by a wide margin over both an LLM layout planner and a specialized layout transformer, while the match-and-place stage bridges the remaining fidelity gap to the ground-truth composites.
Abstract
Automating graphic design synthesis from user-provided elements requires both a coherent overall composition and the exact preservation of each asset. Existing methods predict a layout as explicit bounding-box coordinates with a language model and then paste the assets into it, which separates spatial planning from visual synthesis and tends to produce rigid, mis-scaled compositions. We instead ask whether the layout can emerge implicitly inside a pretrained image-editing diffusion transformer. We present Mise-en-Sc\`ene, a two-stage framework. In the first stage, a diffusion transformer adapted with a small, knockout-selected LoRA drafts a complete design in which the arrangement of the elements emerges jointly with the rendered canvas. In the second stage, a deterministic match-and-place step moves the original high-resolution assets to the drafted positions, which guarantees exact asset fidelity and yields an editable, layered design that a designer can keep refining rather than a flat image. Notably, a minimal adaptation of the pretrained transformer already suffices, without the specialized conditioning machinery commonly introduced for multi-element generation. On the large-scale PrismLayersPlus benchmark, the designs produced by Mise-en-Sc\`ene are the closest to the ground truth in perceived quality among all compared methods, by a wide margin over both an LLM layout planner and a specialized layout transformer, while our match-and-place stage bridges the remaining fidelity gap to the ground-truth composites.
GRE-Diff, a controllable and interactive diffusion-based framework that automates the creation and editing of apartment floor plans under user-specified constraints, is proposed, offering a practical step toward bridging AI-driven automation and human creativity in spatial design.
Jing Wang, Haoran Xiong, Zihao Yan et al.· 0 citations
This work presents StructuredEdit, a pipeline that reframes design editing as parameter manipulation rather than pixel generation and embeds hard design constraints into vision-language model fine-tuning by backpropagating pixel-level constraint violations through a lightweight differentiable rasterizer.
We introduce UniWorld-Design, a framework that redefines image generation from flat pixel synthesis to structured visual composition, with semantic RGBA layers as the atomic units of generation, understanding, and editing. Our key insight is that pixels define how an image is rendered, whereas layers define how an imag...
Zongjian Li, Zhi-Yuan Yan, Chenxu Bai et al.· 0 citations
A post-training paradigm that integrates multimodal alignment (MA) and structural perception (SP) is proposed, MA enhances element interpretation by grounding metadata-defined elements to their visual counterparts, mitigating semantic drift, and SP models layer-aware inter-element spatial relationships to improve hiera...
Yiyang Huang, Zhao-Wen Wang, Simon Jenni et al.· 0 citations
Experiments demonstrate that the MultiCube method can generate high-quality compositional 3D objects with precise part-level control, including those with unique layouts difficult to achieve with text or image prompting alone.
Ava Pun, Kang-Le Deng, Yi-Heng Zhu et al.· 0 citations
The goal is not to present AI as a replacement for artists, but to show how controllable AI systems can support more precise, collaborative, and extensible forms of creative production.
Zheng Wei, Yuying Tang, Mia Tang et al.· Proceedings of the Special I...· 0 citations