Skip to content
Open access

A Fusion-based Machine Learning Approach for Classification of Seabed Defects

Jul 2026 · Earth Systems and Environment · 0 citations · 18 references

TL;DR

The IEResViT model, a novel fusion-based hybrid architecture that integrates a Depthwise Inception Convolutional Neural Network (CNN) with a modified Vision Transformer (ViT) framework, outperforming several state-of-the-art deep learning models while using fewer parameters and reduced computational cost.

Abstract

Accurate classification of the seabed defects is crucial in safeguarding the safety, reliability, and maintenance of underwater infrastructure, including pipelines, cables, and offshore platforms. Classification of underwater images remains a challenging task due to issues such as low visibility, scattering, color distortion, and multi-scale variations in underwater environments. To address these challenges, this study proposes IEResViT, a novel fusion-based hybrid architecture that integrates a Depthwise Inception Convolutional Neural Network (CNN) with a modified Vision Transformer (ViT) framework. The proposed model replaces the traditional multilayer perceptron (MLP) block in the transformer encoder with a ResNet-based residual connection (ResMLP) to enhance hierarchical feature extraction and reduce computational complexity. The architecture processes images through parallel CNN and transformer branches to capture both local and global feature representations, followed by feature fusion and classification. The proposed model was evaluated on the Aquatic Defect dataset, AQUA20 dataset and Marine_PULSE dataset, achieving high classification performance with an accuracy of 98.88% 98.65% and 95.65% respectively, outperforming several state-of-the-art deep learning models while using fewer parameters and reduced computational cost. The results demonstrate that the IEResViT model provides an efficient and lightweight solution for reliable underwater defect image classification. Underwater infrastructure, such as pipelines, communication cables, offshore wind farms, and marine habitats, is a crucial interface on the seabed. Natural processes and human activities may cause seabed defects, e.g., cracks and erosion scars, or sediment liquefaction areas or uncovered buried utilities, over time. Early and precise diagnosis of these defects is critical in preventive maintenance, environmental safety, and navigation. Traditional seabed relies on manual interpretation, and its time consuming and labour-intensive. Recent advancements in machine learning help in automated detection. The graphical abstract presents a deep learning model to classify the aquatic defect images. This article presents the IEResViT model, which uses two parallel branches: a depthwise inception CNN branch that extracts local spatial features, and a ResMLP ViT branch that captures global dependencies through multi-head self-attention mechanisms. The extracted local and global features are then fused together, enabling enhanced representation of complex underwater defect image characteristics. The fused features are subsequently passed to fully connected dense layers, followed by a softmax classifier to predict the defect category. The Aquatic defect image dataset, AQUA20 dataset, and Marine_PULSE dataset are used to evaluate the performance of the model, and achieved an accuracy rate of 98.88%, 98.65%, and 95.65% respectively.

Read PDF

Similar papers

Open access Jul 2026

A Method for Detecting Multiple Types of Defects in Concrete Dams Based on an Improved YOLOv12 Model

Accurate detection and characterization of surface defects in concrete dams is vital for ensuring safe operation. To address the limitations of existing research focused solely on cracks and the challenges traditional convolutional networks face in adapting to deformation and multiscale features, this study introduces DCN-YOLO, a deformable convolution-augmented framework for the simultaneous detection and classification of multiple defect types from UAV-acquired imagery. The model outputs bounding box localizations and categorical labels. Based on YOLOv12, this proposed model integrates DCNv4 deformable convolutions with the C3k2 module. By leveraging adaptive sampling offsets and dynamic modulation, the proposed model enhances geometric modeling for irregular defects, improving the detection of small and medium defects while achieving an acceptable trade-off in inference efficiency. To address multiple defect coexistence, we adopt Binary Cross-Entropy (BCE) loss to decouple classification and localization, improving training stability in multi-label scenarios. A Multi-defects dataset was created using UAV images, and performance was validated on the CrackSeg public dataset. The proposed model achieved 77.4% ± 0.2% overall precision under complex conditions, exceeding the YOLOv12l baseline by 7.1% and improving mAP50-95 by 4.2%. It demonstrated competitive performance in detecting cracks, aggregate exposure, and construction joints, thereby providing a potentially robust and efficient approach for intelligent inspection of concrete dam surface defects.

Wenhao Xu, Wenjie Zhang, Bo Xu · 0 citations
Open access 2026

A Hybrid Multi-Backbone Learning Framework for Marine Pollution Detection Using Deep Feature Fusion

Marine ecosystems are increasingly threatened by increasing industrial activities and anthropogenic environmental impacts. Especially oil spills and chemical wastes cause serious pollution in the seas. This situation negatively affects both marine life and economic activities. In this context, automatic and early detection of marine pollution is of great importance for the effectiveness of environmental response processes. In this study, a hybrid approach combining Transformer-based feature extraction and deep neural networks is proposed for image-based marine pollution classification. Advanced visual feature extraction models such as BEiT, ViT, DeiT, SwinV2, ConvNeXt and DTFC-Net are used and the features extracted from these models are trained with a deep neural network classifier. Model performance is evaluated on two different public datasets using both 5-fold cross-validation and hold-out method. According to the results, all models achieved high accuracy values; however, the proposed DTFC-Net model stood out by achieving the highest accuracy (98.60%, 98.33%) on both datasets. These findings provide evidence that transformer-based deep learning approaches can be effectively applied to marine pollution detection and demonstrate promising performance in environmental monitoring tasks.

Yasin Ozkan · 0 citations
Jul 2026

Deep learning-based detection technology for railway track surface defects

Railway tracks are essential to the transportation system and play a crucial role in rail transit operation. However, during long-term operation, defects such as cracks, wear, deformation, and corrosion are prone to occur on the rail surface. These defects may lead to train derailment or equipment failure, thereby causing serious safety accidents. To address these issues, this paper proposes a railway track surface defect recognition method based on an improved YOLOv11 model, aiming to achieve high-precision and real-time detection. In the neck network, a fused weighted bidirectional feature pyramid network is introduced, enabling the model to adaptively learn and adjust feature fusion weights. In the backbone network, a DSConv-C3K2 dynamic snake convolution enhancement module is designed to better capture slender and irregular defect features. Performance evaluation on the dataset and ablation experiments verify that the proposed YOLOv11-WD model significantly improves detection accuracy while achieving a preliminary lightweight design compared with the baseline YOLOv11s. In addition, the improved modules demonstrate a clear synergistic effect. For the recognition of four types of railway track defects, the baseline YOLOv11s model achieves an mAP@0.5 of 93.5%, while the final improved model reaches 96.4%, marking a 2.9 percentage point improvement over the original YOLOv11s.

Xianwei Zhang, Botong Song · 0 citations
Conference Jul 2026

Application and enhancement of deep learning in infrared image recognition of marine surface vessel targets

In this study, we proposed an improved model based on YOLOv8n network, which aims to the issue of high false alarm rate caused by marine surface clutter in the recognition task of infrared image of vessel target. By integrating StarNet--a backbone network that based on “star operation”, and adding deformable attention mechanism (DAT), our model improved its ability in separating noise from structured signal, and realized data-driven adaptive focus. Compared to the original, our Experiment based on public dataset show that the model our putted up reduces false alarm rate by 2%. At the same time, this achievement comes with a reduction computational load and numbers of parameters—by 17.5% and 21.0%. we convinced it achieved a balance between robustness and lightweight design. This study provides an effective method for the task of detecting infrared vessel targets under complex clutter backgrounds. Moreover, its lightweight design is beneficial to the practical deployment on devices that lack of computing power.

Xinyu Liu · 0 citations
Aug 2026

Automated damage detection and assessment for underwater bridge structures using multi-scale image fusion enhancement and deep learning

The integrity and safety of underwater bridge structures can be compromised by damage; therefore, timely detection and assessment are crucial. However, underwater damage detection is constrained by turbidity, low illumination, and multiple coexisting damage types, which complicates comprehensive automated safety assessment. This study proposes an automated framework for structural damage detection and safety assessment of underwater bridge structures. Multiple types of underwater damage are analyzed, and an underwater damage image dataset (UDID) is established. A modified linear unsharp masking method is used to adaptively enhance the high-frequency features of the UDID through multi-scale image fusion. A well-trained GoogLeNet model is used to automatically detect underwater structural damage. Based on the detection results, bridge safety is classified into five levels using the damage index method. A case study involving a concrete bridge in China demonstrates the effectiveness of the proposed framework. The GoogLeNet model achieves a detection accuracy of 97% in the large-scale test. The concrete bridge is assessed as level 3, which represents moderate damage and is consistent with field detection results. In underwater environments featured by a low signal-to-noise ratio, the detection accuracy of damage types that depend on texture and contrast features decreases significantly. This framework effectively addresses the limitations of existing underwater damage detection methods and enables automated safety assessment for underwater bridge structures.

Yang Deng, Xueqi Cao, Jialiang Guo et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.