This work developed Mask-based Weighted Conditional Flow Matching (MWCFM), which extends Flow Matching by introducing contextual masks for precise feature focusing, which enables targeted training on critical spatial elements relevant to urban planning.
Abstract
In the context of urban planning, architects are normally instructed with creating presentation images that visualize proposed buildings within their urban context. This work aims to develop a GenAI model for automatically generating architectural presentation images in urban scenes, with emphasis on model optimization. To achieve this, we developed Mask-based Weighted Conditional Flow Matching (MWCFM), which extends Flow Matching by introducing contextual masks for precise feature focusing. This enables targeted training on critical spatial elements relevant to urban planning. Our trained model learns from urban street-view data while adhering to specific style-guidelines, which are integrated into training through the loss function. Furthermore, the model's performance is evaluated using application related metrics, derived from presentation image style guidelines.
A perception-aware framework for green streetscape redesign that transforms a real street-view image and a high-level design goal into a realistic visualization of a retrofitted street, and suggests that controllable generative models can provide practical support for urban street retrofit visualization and assessment.
Hongkun Wang, Fei Li· Digital Signal and Computer...· 0 citations
Streetscape quality has become a central concern in contemporary urban planning, particularly within the framework of the pedestrian-friendly 15-minute city, where walkability and public-space quality are increasingly recognized as key determinants of urban performance. However, assessing streetscape qualities across large suburban and peri-urban territories remains challenging due to the time and resource demands of conventional field surveys. This paper presents a planning-oriented assessment of streetscape qualities in the north-eastern periphery of Nice (France) using the latest release of SAGAI (Streetscape Analysis with Generative AI), an open-source workflow that leverages vision-language models (VLMs) for large-scale streetscape analysis from Google Street View imagery. The new release addresses limitations of the original framework through improved image acquisition, geographically consistent view generation, support for multiple VLM architectures, consensus-based inference, and an integrated analytical environment. The workflow is applied to several thousand street-level observations to evaluate qualities relevant to pedestrian-friendly urban environments: sidewalk presence, pedestrian entrance density, and vegetation. The resulting maps reveal that the desired streetscape qualities characterize only a fraction of today's suburban streetscapes, mainly in compact developments and traditional suburban faubourgs, while they are particularly lacking on residential hills. The analysis demonstrates the potential of contemporary VLMs to support urban diagnostics in extensive suburban territories where fieldwork would be prohibitively time-consuming. Beyond the case study, the paper illustrates how recent advances in vision-language models can contribute to evidence-based planning by enabling scalable, flexible, and interpretable assessments of urban public-space quality.
Building materials and their spatial distribution play a significant role in determining the outcomes of human-caused and natural disasters in urban and peri-urban areas. However, building-level data on building and roofing materials are scarce. Here, we explore the feasibility and performance of a Convolutional Neural Network (CNN) model using spectrally transformed high-resolution multispectral imagery to map roofprints (i.e., classifying and delineating roofing materials at the building footprint-level) in Washington, District of Columbia (D.C.) and Denver, CO, United States. To generate consistent training data, we integrate geospatial vector data of individual building footprints with real estate industry-derived building-level roofing material data to create labeled image data from Planet SuperDove imagery. We compare the CNN classifier to a pixel-based machine learning (ML) model to demonstrate the capability of our roofprints mapping approach. With F1-scores ranging from 0.56 to 0.95 for the most common roof material classes, the CNN model outperformed the pixel-based ML classifier by 15% and 17% in Washington, D.C., and Denver, respectively. Results demonstrate within-domain robustness for the studied metropolitan areas, which are characterized by differing building densities, roof morphologies, and material patterns. While cross-region transferability was not evaluated, our findings provide a controlled comparison of pixel-based and context-aware approaches for rooftop material mapping and highlight the importance of hierarchical representations that integrate spectral information with roof texture, edge characteristics, spatial arrangement, and neighborhood context for improving classification performance. Accurately mapping building materials has the potential to advance urban planning and environmental policies, including assessments of heat exposure, energy demand, as well as hazard risk and community resilience.
C. Amaral, Maxwell C. Cook, Johannes H. Uhl et al.· Remote Sensing· 0 citations
Abstract. Semantically rich 3D city models play a vital role in a variety of applications, such as urban planning. Enhancing these models with currently unavailable attributes, such as building storey numbers, can unlock new opportunities to address pressing challenges, including sustainable urban development. In this work, we present an end-to-end pipeline for the automatic estimation of the number of storeys to semantically enrich 3D city models. We employ volunteered geographic information street-view imagery from Mapillary, using a COCO-pretrained object detection model to identify windows in fac¸ade images as key visual indicators for inferring building storey counts. Our detection pipeline, based on the YOLOv3 architecture, estimates storey numbers using an ensemble of clustering methods including Gaussian Mixtures and DBSCAN and enables the automatic augmentation of CityGMLbased 3D city models by filling in missing attributes. This enrichment supports advanced applications, such as assessing buildingscale energy demand, evaluating vertical urban growth patterns or population density estimations. We validated the feasibility of our approach with unfiltered Mapillary and applied it to a district in the city of Heidelberg, Germany. The paper also includes a detailed discussion of learning process quality, integration workflows, and visualization of the enriched 3D city model. The developed code is available at: https://github.com/hcu-cml/citydb-buildingstoreys-ai.
Lukas Arzoumanidis, Al Maimun As Samee, Elmehdi Kanna et al.· ISPRS Annals of the Photogra...· 0 citations
Accurate vector mapping of buildings and walls is critical for geospatial applications but remains a labor-intensive process. While recent deep learning methods have improved automatic extraction, in order to meet cartographic standards they always require a human to perform quality control and fix complex cases in the extraction. We present Click2Poly, a human-in-the-loop AI assistant designed to speed up this manual step. Extending the Florence-2 Vision Language Model (VLM), Click2Poly responds to user clicks by editing the building or wall vector layer directly. Implemented as a QGIS plugin, Click2Poly speeds up the manual editing of building and wall vector layers in a real-world production environment.
Nicolas Girard, Jawher Ben Abdallah, Arno Gobbin et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.