Skip to content
#generative ai Open access

A Comparative Analysis of Generative AI for Architectural Visualization: A Quantitative Comparison Using Objective Evaluation Metrics and the Architectural AI Score (AIS)

Oct 2026 · Applied Sciences · 10 references
Generative Adversarial Networks and Image Synthesis

Abstract

This study presents a quantitative metric-based comparison of four state-of-the-art text-to-image generative AI systems for architectural visualization: OpenAI Image Generation, Imagen 4, Midjourney v8.1, and Stable Image Ultra. The models were evaluated using the Architectural Prompt Dataset (APD-50), a controlled, expert-informed prompt set consisting of 50 architectural prompts across five typologies: Modern Residential Buildings, Public Buildings, Religious and Cultural Architecture, Interior Design, and Facade and Sustainable Architecture. For each prompt, four images were generated per model–prompt combination, resulting in a corpus of 800 images. The generated images were analyzed using reference-free and image-level computer vision metrics, including CLIP Score for prompt alignment, Raw BRISQUE and Quality_BRISQUE for reference-free perceptual quality, Shannon Entropy as a diagnostic descriptor of visual information density, LPIPS as intra-prompt perceptual variation, and SSIM as intra-prompt structural similarity. In response to the limitations of treating entropy as an inherently positive indicator, Shannon Entropy was excluded from the revised composite score and retained only as a diagnostic metric. Alongside the individual metric results, this study proposes the Revised Architectural AI Score (AIS-R) as a first-stage, exploratory composite index for summarizing selected visual-output characteristics under controlled prompt conditions. AIS-R combines normalized CLIP, Quality_BRISQUE, and LPIPS components using an explicit and reproducible weighting structure. The purpose of AIS-R is not to replace professional architectural judgment or to establish a fully validated architectural-quality framework at this stage. Rather, it provides an initial operational formulation for comparing prompt alignment, reference-free perceptual quality, and intra-prompt perceptual variation within the limited scope of the APD-50 experimental setup. Linear mixed-effects models with prompt-level random intercepts were used as the primary inferential framework. Under the revised metric configuration, Midjourney v8.1 achieved the highest AIS-R score (0.6504 ± 0.0586), followed by Stable Image Ultra (0.6267 ± 0.0582), Imagen 4 (0.6069 ± 0.0707), and OpenAI Image Generation (0.5829 ± 0.0555). However, these differences should be interpreted as metric-based differences in selected visual-output characteristics rather than as evidence of architectural quality, constructability, spatial validity, scale accuracy, or professional workflow suitability. The study therefore represents an initial metric-formulation and controlled benchmarking stage, while future research should examine the correspondence between AIS-R and professional architectural judgments through expert evaluation, user studies, or workflow-based validation.

View source

Similar papers

#artificial intelligence Conference Open access Apr 2020

ECCOLA - a Method for Implementing Ethically Aligned AI Systems

The method, ECCOLA, is presented, which aims at making the high-level AI ethics principles more practical, making it possible for developers to more easily implement them in practice.

Ville Vakkuri, Kai-Kristian Kemell, P. Abrahamsson · 64 citations · ⚡6
#computer vision Review Apr 2024

AI-powered Code Review with LLMs: Early Results

The goal is to not only refine the accuracy of the LLM-based tool but also to underscore its potential in streamlining the software development lifecycle through proactive code improvement and education.

Z. Rasheed, Malik Abdul Sami, Muhammad Waseem et al. · 62 citations · ⚡3
#computer vision Open access Mar 2024

LLM-based agents for automating the enhancement of user story quality: An early report

The use of large language models to automatically improve the user story quality in Austrian Post Group IT agile teams is explored, with a reference model for an Autonomous LLM-based Agent System developed and implemented at the company.

Zheying Zhang, M. Rayhan, Tomas Herda et al. · 48 citations · ⚡4
#computer vision Review Mar 2024

System for systematic literature review using multiple AI agents: Concept and an empirical evaluation

This paper introduces a novel multi-AI-agent system designed to fully automate SLRs, and demonstrates how it substantially reduces the time and effort traditionally required for SLRs while maintaining comprehensiveness and precision.

Abdul Malik Sami, Z. Rasheed, Kai-Kristian Kemell et al. · 44 citations · ⚡2
#computer vision Feb 2024

Can Large Language Models Serve as Data Analysts? A Multi-Agent Assisted Approach for Qualitative Data Analysis

The proposed LLM-based multi-agent system automates qualitative data analysis process, creating opportunities for researchers and practitioners, and future improvements focus on enhancing multilingual performance and integrating continuous expert feedback.

Z. Rasheed, Muhammad Waseem, Aakash Ahmad et al. · 41 citations

Related blog posts

Microsoft Research Blog Oct 7, 2026

Agent Lightning v1.0: A 3,500-Line Lightweight Agentic RL Framework for Training Agents with Real Harnesses

Training AI agents with reinforcement learning can be challenging because their tools, context, and decision-making are managed by complex frameworks. Agent Lightning connects existing agents to RL training, making it easier to improve them without rebuilding them. The post Agent Lightning v1.0: A 3,500-Line Lightweight Agentic RL Framework for Training Agents with Real Harnesses appeared first on Microsoft Research.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.