Jul 2026· Digital Signal and Computer Communications· Vol 14294, pp. 142941I - 142941I-6· 0 citations· 16 references
Engineering
TL;DR
A perception-aware framework for green streetscape redesign that transforms a real street-view image and a high-level design goal into a realistic visualization of a retrofitted street, and suggests that controllable generative models can provide practical support for urban street retrofit visualization and assessment.
Abstract
Urban street retrofitting is increasingly used to improve greenery, walkability, and perceived safety, yet planners and communities often lack intuitive visualizations of how a street may look after such interventions. This paper presents a perception-aware framework for green streetscape redesign that transforms a real street-view image and a high-level design goal into a realistic visualization of a retrofitted street. The proposed framework integrates an MLLM-guided planner for structured redesign operations, a rule-based compiler for semantic mask editing, a ControlNet-guided diffusion renderer for candidate generation, and a perception-aware selector for choosing the final design. Experiments on Cityscapes and Mapillary Vistas, show that the method achieves a favorable balance among environmental improvement, image realism, and structural preservation. Additional comparisons further demonstrate the value of multimodal planning and perception-aware ranking. These results suggest that controllable generative models can provide practical support for urban street retrofit visualization and assessment.
Streetscape quality has become a central concern in contemporary urban planning, particularly within the framework of the pedestrian-friendly 15-minute city, where walkability and public-space quality are increasingly recognized as key determinants of urban performance. However, assessing streetscape qualities across large suburban and peri-urban territories remains challenging due to the time and resource demands of conventional field surveys. This paper presents a planning-oriented assessment of streetscape qualities in the north-eastern periphery of Nice (France) using the latest release of SAGAI (Streetscape Analysis with Generative AI), an open-source workflow that leverages vision-language models (VLMs) for large-scale streetscape analysis from Google Street View imagery. The new release addresses limitations of the original framework through improved image acquisition, geographically consistent view generation, support for multiple VLM architectures, consensus-based inference, and an integrated analytical environment. The workflow is applied to several thousand street-level observations to evaluate qualities relevant to pedestrian-friendly urban environments: sidewalk presence, pedestrian entrance density, and vegetation. The resulting maps reveal that the desired streetscape qualities characterize only a fraction of today's suburban streetscapes, mainly in compact developments and traditional suburban faubourgs, while they are particularly lacking on residential hills. The analysis demonstrates the potential of contemporary VLMs to support urban diagnostics in extensive suburban territories where fieldwork would be prohibitively time-consuming. Beyond the case study, the paper illustrates how recent advances in vision-language models can contribute to evidence-based planning by enabling scalable, flexible, and interpretable assessments of urban public-space quality.
This work developed Mask-based Weighted Conditional Flow Matching (MWCFM), which extends Flow Matching by introducing contextual masks for precise feature focusing, which enables targeted training on critical spatial elements relevant to urban planning.
Katharina Roth, Eva Hagen, Alexander Bartscher et al.· 0 citations
Understanding urban perception from street‐view imagery has become a central topic in urban analytics and human‐centered urban design. However, most existing studies treat urban scenes as static and largely ignore the role of dynamic elements such as pedestrians and vehicles, raising concerns about potential bias in perception‐based urban analysis. To address this issue, we propose a controlled framework that isolates the perceptual effects of dynamic elements by constructing paired street‐view images with and without pedestrians and vehicles using semantic segmentation and generative inpainting powered by large vision models. Based on 720 paired images from Dongguan, China, a perception experiment was conducted in which participants evaluated original and edited scenes (removal of pedestrians and vehicles) across six perceptual dimensions, including wealth, safety, vibrancy, beauty, boredom, and depression. The results indicate that removing dynamic elements leads to a consistent decrease in perceived vibrancy (30.97%), whereas changes in other dimensions are more moderate and heterogeneous. To further explore the underlying mechanisms, we trained 11 machine‐learning models using multimodal visual features. Feature‐importance analysis identifies lighting conditions, human presence, and depth variation as key factors driving perceptual change. At the individual level, 65% of participants exhibited significant vibrancy changes, compared with 35%–50% for other dimensions, revealing notable inter‐individual sensitivity; gender further showed a marginal moderating effect on safety perception. Beyond controlled experiments, the trained model was extended to a city‐scale dataset covering 47,963 locations (191,852 images) to predict vibrancy changes after the removal of dynamic elements. The city‐level results reveal that such perceptual changes are widespread and spatially structured, affecting 73.7% of locations and 32.1% of images, suggesting that urban perception assessments based solely on static imagery may substantially underestimate urban liveliness. Overall, this study highlights the critical role of dynamic elements in shaping urban perception and underscores the importance of accounting for transient urban features in large‐scale perception‐driven urban studies.
Zhiwei Wei, Meng-Zi Zhang, Boyan Lu et al.· Transactions on GIS· 0 citations
Abstract. Urban digital twin systems require 3D city representations that reconcile semantic structure, geometric reliability, simulation capability, and photorealistic real-time rendering. Existing approaches usually prioritize a single modelling paradigm, limiting their ability to support both analytical and visualization needs. CityGML provides standardized semantics and topology but often lacks surface realism. Semantic mesh models preserve geometric detail suitable for environmental simulations but provide limited hierarchical semantics. In contrast, neural radiance-field approaches such as 3D Gaussian Splatting (3DGS) enable photorealistic rendering at interactive frame rates but do not explicitly encode topology or structured semantics. This study establishes a comparative framework linking LiDAR-derived CityGML, semantic mesh, 3D Gaussian Splatting, and Triangle Splatting within a unified urban modelling workflow. UAV data acquired using a DJI ZENMUSE L2 sensor serve as the geometric backbone for reconstructing CityGML LoD1–LoD2 models. The semantic model is transformed into a textured triangular mesh, while radiance-based models are generated from the same imagery using multiple 3DGS implementations and a triangle splatting framework. Comparative evaluation investigates geometric coherence, semantic preservation, and radiance consistency to identify structural correspondences across the representations. The results reveal complementary modelling layers that can be systematically mapped rather than treated as competing alternatives. Based on these findings, the paper proposes a conceptual foundation for a unified 3D urban model capable of transforming consistently into semantic, surface-based, and radiance-based representations for adaptive urban digital twin systems. Data are freely accessible for research purposes at https://github.com/3DOM-FBK/urban-representation-fusion/.
D. Suwardhi, Muhammad Arif Sudibyo, Agus Ambarwari et al.· The International Archives o...· 0 citations
A human-in-the-loop GenAI-assisted framework for producing immersive 3D visualization prototypes rather than historically verified reconstructions is proposed, which integrates multi-view image generation, knowledge-informed review, single-image-to-3D generation, topology inspection, and perceptual calibration.