Jul 2026· 2026 7th International Conference on Smart Systems and Inventive Technology (ICSSIT)· pp. 2137-2142· 0 citations· 18 references
Abstract
Street View Images (SVI) are high-resolution, geo-referenced panoramas that capture real-world environments. Integration of Artificial Intelligence (AI) with SVI enables automated analysis for a range of urban applications including object detection, semantic segmentation, text recognition, scene understanding, and socioeconomic prediction. More than 25 recent AI based SVI studies covering applications in crime prediction, building attribute classification, sidewalk inventory, land price estimation, and environmental monitoring were screened for the survey. Across domains, deep learning architectures such as ResNet, ConvNeXt, and Vision Transformers consistently outperformed traditional machine learning models, with reported accuracies up to 94% for classification and R2 values between 0.62–0.83 for prediction tasks. SVI augmented with other data sources like satellite imagery and Global Information System (GIS), enhanced the model performance and contextual understanding. The Findings highlight the dominance of convolutional and transformer-based networks, emerging interest in graph neural networks, and the need for generalized models and diverse datasets to advance SVI research.
As remote sensing technologies improve, we are now able to look at the Earth from different points of view. These technologies have enabled major changes in many areas. High-resolution satellite images have enabled progress in civilian and military applications such as environmental monitoring, disaster management, and public safety. While these images give us much information, we also need advanced analysis and automated object detection. Recent studies employing deep learning architectures reached high success rates in object detection from satellite imagery. This study introduces a novel model by assessing the impacts of several YOLO object detection algorithms with the Convolutional Block Attention Module (CBAM) on aircraft detection from satellite images. We used the HRPlanesv2 dataset for our experiments. The results revealed that the proposed model had better performance compared to alternative models. The proposed model achieved a mAP50 of 0.9868, a mAP50-95 of 0.7911, a precision of 0.9811, and a recall of 0.9637. Incorporating CBAM improves the detection of objects in crowded and complex scenes. These results demonstrate that attention mechanisms have a significant impact when used with the YOLO architecture for object detection in satellite images. It also provides a reliable and efficient solution for practical applications requiring accurate and consistent aircraft detection.
Ibrahim Aruk, Hakan Açıkgöz, Ertuğrul Doğruluk· Konya Journal of Engineering...· 0 citations
Street space facilitates public life within communities, and advancements in artificial intelligence technology bring innovative approaches to assessing and forecasting street space perceptions. This study is centered on the Tsim Sha Tsui area of Hong Kong, as an illustrative case. Street view images and historical photographs are used as primary data sources, leveraging deep learning for semantic segmentation to measure the constituent elements of street space. A perceptual evaluation model is developed using regression analysis and time-series forecasting models to quantify and describe in detail the regional styles and their evolution throughout the historical development of urban streets. The findings indicate that architectural elements consistently account for an average of 42.3% historically, while the proportion of carriageways has significantly increased since the 1980s, rising by 1.8% annually. Although the rate of greening has risen, sky visibility has declined, leading to a decreased level of openness. Utilizing ordinary least squares regression in conjunction with long short-term memory time-series predictions, it is anticipated that the share of buildings and signage will reach 57.6% by 2050 (+15.3% compared with 2020), with the share of sky and greenery decreasing to 18.9%. Subsequently, by integrating the quantitative results and substituting relevant assumptions, projections regarding the future transformations of street styles are made. Artificial intelligence–generated content (AIGC) technology is applied to adjust weights, enabling the fusion and transition of various elemental styles. The Stable Diffusion model is employed to generate three images representing future styles (traffic city, green city, and consumer city), and the correlation between the quantitative indicators and style migration is validated. This study demonstrates that the methodology employed can effectively analyze the historical evolution of street space and forecast multiple future scenarios, thereby providing data support for the revitalization of high-density urban streets. On the one hand, it can inform planning decisions by establishing thresholds for elemental occupancy, and on the other hand, it can assist in the selection of various scenarios through stylized image generation. Implementing street space quantification techniques and forecasting future styles provide valuable support for designers and evaluators of urban street design schemes, providing crucial insights for the design of street space and architectural form.
Tian-Lian Wang· Journal of urban planning an...· 0 citations
This paper addresses the semantic segmentation of road images, a critical task for autonomous vehicle navigation, particularly in non-urban environments that present significant challenges. While much research focuses on well-maintained roads in developed countries, this study confronts the complexities of real-world conditions, such as those prevalent in developing nations, which feature vast networks of unpaved and poorly maintained roads. The core of our methodology is a neural network architecture that synergistically combines the encoder-decoder structure of U-Net with the feature extraction power of a ResNet backbone. The primary objective is the precise classification of each image pixel into one of four essential categories for navigation: background, asphalt, paved, and unpaved road. The model's training regimen involved exploring different ResNet versions (ResNet18, ResNet34, and ResNet50) as the encoder backbone to assess the impact of network depth. A key aspect of our approach was a progressive training strategy, where model versions were trained on images of varying resolutions. The results demonstrated a significant and somewhat counter-intuitive finding: training the ResNet34-U-Net model with images at half the original resolution yielded the best overall performance, achieving the highest Dice and IoU scores. This suggests that reducing image resolution acts as an effective form of regularization, compelling the model to learn more general and robust features by ignoring minor, irrelevant details. This outcome not only enhances the model's generalization capabilities for diverse and imperfect road conditions but also carries a substantial practical advantage by reducing the computational cost of training and inference, a crucial factor for deployment on resource-constrained embedded systems in autonomous vehicles.
Victor M. A. do Nascimento¹, André T. Cunha Lima1, Nascimento. Av· JOURNAL OF BIOENGINEERING, T...· 0 citations
SRB-Net is presented, a U-Net-based framework that combines three complementary components: strip pooling for long-range horizontal and vertical context; residual multi-scale atrous spatial pyramid pooling with squeeze-and-excitation blocks for multi-scale and channel-aware feature learning; and a bottleneck attention module (BAM) for refining skip-connection features.
Hamdoun Youssef, Xingyuan Li, Yongtao Yu et al.· Journal of Computing and Ele...· 0 citations