Aug 2026· Remote Sensing· Vol 18, pp. 2884· 0 citations· 22 references
Abstract
The identification of visible submerged hazards, including shallow-water shipwrecks and associated debris, is important for maritime safety, coastal management, environmental monitoring, and marine archeology. The scope is restricted to wrecks that remain optically visible from above in shallow or intertidal water. This study presents Submerged Hazard Identification and Processing using Augmented Image-based Detection (SHIP-AID), a modular GeoAI evaluation framework for high-resolution RGB imagery. Following site-level quality control, the independent source dataset contains 695 images from 403 wreck sites. The dedicated group-disjoint holdout contains 150 images from 88 sites and 184 annotated wreck objects. Five detector backbones and one task-aware underwater-enhancement baseline were evaluated using ten matched training seeds. Under the standard benchmark evaluation protocol, in which precision and recall are reported at the internally determined maximum-F1 point of the confidence sweep, the best configuration achieved precision 0.896, recall 0.861, mAP@50 0.927, and mAP@50–95 0.668 on the locked holdout. At the fixed, validation-selected operating threshold of 0.45, the primary detector produced threshold-specific precision 0.922 and recall 0.837. Site-clustered bootstrap intervals were 0.895–0.951 for mAP@50 and 0.625–0.704 for mAP@50–95. Physically informed attenuation and backscatter augmentation improved stricter-IoU performance relative to generic augmentation, whereas global Otsu thresholding reduced recall and localization accuracy. Performance remained stable under mild degradation, declined under moderate and strong degradation, and became unreliable under severe low visibility. SHIP-AID is therefore positioned as a decision-support framework for prioritizing optically visible shallow-water sites, with sonar, diving, hydrographic, or archeological evidence retained as the confirmation standard.
To address the accuracy degradation of ship detection caused by occlusion from adjacent vessels, shore-based facilities and meteorological obscuration in complex maritime-surveillance scenes, this paper proposes an occlusion-robust detection model named OAR-YOLO. An adaptive dual-path downsampling module termed ADown was embedded at the three backbone levels P3, P4 and P5, in which low-frequency contextual information and high-frequency edge information were preserved separately through parallel average-pooling and max-pooling branches, alleviating the information loss caused by conventional strided-convolution downsampling. An attention-driven intra-scale feature interaction module termed AIFI was embedded at the top level P5 to establish semantic associations between spatially separated visible regions through global self-attention, compensating for the insufficient cross-region connectivity caused by the locality of convolution. The two modules formed a dual compensation mechanism of information conservation and semantic connectivity. On a self-built ship dataset, OAR-YOLO achieved a Precision of 81.0%, an mAP@0.5 of 74.7% and an mAP@0.5–0.95 of 47.1%, with gains of 2.7, 2.6 and 1.4 percentage points over the YOLO11n baseline. The model has only 2.89 M parameters and 5.7 GFLOPs, with an inference time of 0.8 ms per frame, meeting the real-time deployment requirements of complex maritime applications.
Jia-Hang Li, Yan Zhang, Yu Sun et al.· Journal of Marine Science an...· 0 citations
Marine debris threatens aquatic habitats and complicates inspections in ports, seabed environments, and offshore infrastructure. Previous lightweight detectors remain vulnerable to weak texture, blurred boundaries, and cluttered multi-scale features, while direct network expansion conflicts with restricted onboard resources. This study adapts YOLO11n by integrating Receptive-Field Aggregation (RFA) with a Residual Channel-Spatial Recalibration (RCSA) implementation based on dynamic residual groups. Experiments used a public 15-class dataset with 10,884 training images and 1001 model-selection validation images containing 1892 annotated objects. All principal checkpoints were trained for 100 epochs. Across seeds 42, 2026, and 3407, RFA + RCSA achieved validation precision 0.861 ± 0.016, recall 0.801 ± 0.012, mean average precision at IoU 0.5 (mAP@0.5) 0.848 ± 0.001, and mAP@0.5:0.95 0.511 ± 0.002. On an audited group-disjoint holdout (498 images; 963 instances), the corresponding means were 0.820 ± 0.033, 0.748 ± 0.014, 0.779 ± 0.018, and 0.467 ± 0.007. The detector contains 4.19 M parameters and requires 8.91 giga floating-point operations (GFLOPs). These results position it as a lightweight candidate for resource-constrained remotely operated vehicle (ROV) perception; they do not establish real-time embedded deployment.
Object detection in visible nearshore surveillance imagery is of great importance for maritime safety, intelligent coastal monitoring, and water rescue applications. Nevertheless, reliable detection remains difficult because nearshore scenes often contain numerous small targets, cluttered wave patterns, shoreline textures, and substantial illumination variations. Moreover, existing public datasets mainly emphasize vessel detection and provide limited nearshore object categories. To solve these limitations, this study presents a large-scale visible nearshore dataset containing 20,934 images annotated with seven categories: pedestrian, sailor, swimmer, ship, boat, flotage, and seamark. The dataset is designed to support comprehensive evaluation and fair comparison of detection algorithms in complex nearshore environments. Based on the proposed benchmark, we conduct extensive evaluations of multiple mainstream object detectors and further develop a detection framework termed VN-DETR. The proposed model enhances both feature extraction and multi-scale feature fusion for nearshore scenarios. Specifically, a kernel selective attention based on WTConv (WKSA) module is designed to enlarge the receptive field and exploit contextual information in visible images, enabling more accurate object classification. In addition, a cross-layer feature selection and fusion (CFSF) module is introduced to perform feature matching, selection, and fusion across adjacent layers, enhancing the discriminability between foreground objects and complex nearshore backgrounds. This design effectively improves robustness against background noise such as wave reflections and shoreline textures. Extensive experiments on the constructed dataset demonstrate that VN-DETR consistently outperforms representative baseline methods and achieves superior detection performance, particularly for challenging small object categories.
Zhi-Bin Liu, Yong-Jing Jiang, Kao Zhang et al.· Remote Sensing· 0 citations
Abstract. In the initial response to wildfires, securing rapid and accurate geographic information is essential. However, helicopter imagery acquired on-site often lacks precise sensor metadata, such as camera pose and internal parameters, making the application of georeferencing difficult. In particular, obliquely captured wildfire imagery presents additional registration challenges due to severe viewpoint changes, scale variations, and low-texture environments. This study proposes an automated georeferencing pipeline capable of operating under these constraints. The proposed method consists of five stages: preprocessing, image retrieval, feature extraction and matching, Exterior Orientation Parameters (EOP) estimation, and orthomosaic generation. An initial Area of Interest (AOI) is defined using inaccurate initial position data, and the Region of Interest (ROI) within the reference map is obtained through a ResNet50-based image retrieval approach. Subsequently, virtual Ground Control Points (GCPs) are generated through deep learning-based feature matching. Elevation data is then assigned using a Digital Elevation Model (DEM), and EOP are estimated via Perspective-n-Point (PnP) and RANSAC algorithms. Intermediate frames are initialized via interpolation and refined through bundle adjustment to produce the final orthomosaic. Experimental results demonstrated that utilizing SuperGlue and LightGlue complementarily increased the number of successfully georeferenced intervals from 5 to 9. Furthermore, a minimum RMSE of 28.30 m was achieved in the most accurate interval. This method proves that by automating the feature-based georeferencing process, practical geographic information can be rapidly provided for initial disaster response, even in sensor-limited environments.
Seongyun Kim, Jeonghyo Oh, J. Cheon et al.· The International Archives o...· 0 citations
Abstract. Monitoring dynamic alluvial rivers is essential for safe inland navigation, yet traditional bathymetric surveys are costly and infrequent. This paper presents an automated method for detecting migrating sandbars by integrating Sentinel-2 satellite imagery with daily water gauge data. Implemented in Google Earth Engine (GEE), the algorithm matches specific water levels with cloud-optimized images to map emerging shoals. Water and sediment were separated using the Sentinel Water Mask (SWM) index, while a 30-meter internal channel buffer mitigated shoreline mixed-pixel errors. The method’s accuracy was validated using 3-meter resolution PlanetScope imagery. Results demonstrated high geometric agreement (mean Intersection over Union = 0.71) and a strong area correlation (R² = 0.97). Notably, the 10-meter Sentinel-2 resolution caused a systematic 26% overestimation of sandbar size. However, for navigation, this overestimation provides a beneficial safety margin that prevents the underestimation of submerged obstacles. By correlating specific gauge levels with sandbar emergence, the extracted 2D contours provide a vital spatial baseline that enables the future estimation of available water columns over specific bottlenecks. Ultimately, this cost-effective procedure allows for the continuous generation of spatial databases, forming a practical foundation for dynamic relative depth mapping within River Information Services (RIS).
M. Smiarowski· The International Archives o...· 0 citations
Abstract. Effective maritime surveillance and small-scale fisheries management remain challenging in coastal waters, where small vessels are not systematically tracked and are often poorly represented in medium-resolution satellite imagery. Within the AI4COPSEC Horizon Europe framework, this study investigates an object-detection workflow for monitoring small vessels along the Adriatic coasts of Marche and Puglia, Italy. Sentinel-2 and high-resolution PlanetScope RGB imagery were manually annotated to build a task-specific optical dataset and to fine-tune models previously pretrained on a larger SAR-optical vessel dataset. This two-stage strategy was designed to exploit heterogeneous vessel representations during pretraining and then adapt the detector to the target coastal optical domain. The resulting dataset comprised 4,202 image tiles for pretraining and 706 tiles for fine-tuning, with 16,096 and 1,716 vessel annotations, respectively, all belonging to a single target class. Detection experiments were conducted using several YOLOv26 configurations trained under a consistent protocol to assess the trade-off between model complexity, accuracy, and computational efficiency. Among the standard variants, YOLOv26-M achieved the most balanced performance, with Precision of 0.813, Recall of 0.846, F1-score of 0.829, Accuracy of 0.719 and mAP50-95 of 0.306. Pruned and lightweight alternatives showed competitive efficiency-oriented behaviour. Results indicate that, in small-target coastal environments, increasing model size does not necessarily yield proportional gains, whereas task-oriented architectural design improves the balance between detection quality and computational cost.
S. Chiappini, A. Galdelli, Andrea Fiorani et al.· The International Archives o...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.