Beyond Atomic Layouts: Compositional Design Understanding with Vision-Language Models
A post-training paradigm that integrates multimodal alignment (MA) and structural perception (SP) is proposed, MA enhances element interpretation by grounding metadata-defined elements to their visual counterparts, mitigating semantic drift, and SP models layer-aware inter-element spatial relationships to improve hierarchical understanding and reduce structural ambiguity.