Aug 2026· 2026 International Workshop on Intelligent Systems (IWIS)· pp. 1-5· 0 citations· 27 references
Abstract
Cultural heritage preservation increasingly relies on AI-driven semantic segmentation to document and analyze intricately carved stone reliefs. This paper is the first to do a systematic benchmarking study of five deep CNN architectures: DeepLabV3+, PSPNet, SegNet, DenseUNet, and U-Net. We changed these architectures so they could work with four-channel RGB-D input and tested them on a multi-class Borobudur bas-relief dataset with eight semantic categories. All of the models use the same way to train, which is a compound FocalDiceLoss with a clear background exclusion. A multi-dimensional evaluation framework includes the following: segmentation accuracy (mIoU, F1, pixel accuracy), convergence dynamics, parameter efficiency, per-class IoU, qualitative overlay analysis, and pixel-level error mapping. The best mIoU (0.4990), the lowest validation loss (1.1524), the best stability (sigma=0.00068), and the lowest pixel error rate (15.1%) are all found in DeepLabV3+.
Semantic segmentation of the Borobudur temple bas-relief panels is difficult due to stone weathering, class imbalance, complex iconography, and limited annotated data. In this paper, we propose a dual-stream fusion framework, which includes an RGBD branch (RGB+depth, 4-channel) and an Edge-Depth branch (softedge+depth,...
Lukman Awaludin, Wahyono, N. Wirasanti et al.· Engineering, Technology &...· 0 citations
In this paper, we propose RGBD20K, a novel dataset for facilitating the development of more robust and general RGB-D semantic segmentation by encompassing abundant categories and high-quality annotations. RGBD20K possesses several attractive properties: (1) Expanded Semantic Space. In particular, it covers 160 fine-gra...
Shao-Hua Dong, Ze-Xuan Meng, Hai-Yan Sun et al.· 0 citations
In RGB-D semantic segmentation, fusing the RGB branch with the depth-learning branch supports fine-grained semantic interpretation in clutteblue indoor scenes. In current computer vision practice, Transformer architectures are used almost by default, and cross-modal fusion settings follow the same trend. HDBFormer is o...
The results show that model selection in building extraction cannot rest on a single accuracy figure, and that boundary-based metrics should govern the choice in applications where the position of the boundary is critical.
Cafer Yazıcıoğlu, Ammar Aslan· Gazi Üniversitesi Fen Biliml...· 0 citations
To improve the binary semantic segmentation of bare rock and background in complex urban environments, this study developed an improved DeepLabV3+ model for red–green–blue (RGB) imagery derived from Gaofen-1 (GF-1) data. Four backbone networks—Xception, ResNet50, MobileNetV2, and MobileNetV3—were first evaluated. ResNe...
Qiang Wang, Hao-Chuan Lei, Xia-Song Hu et al.· Applied Sciences· 0 citations
This work introduces MonuSegFormer, a heritage-specific hybrid architecture coupling a pretrained Swin-B encoder with a Multi-Scale Atrous Fusion module and a CBAM-augmented progressive decoder to address severe class imbalance in Moroccan historical monuments.
Ouail Choukhairi, Mouad Choukhairi, A. Choukri et al.· Journal of Imaging· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.