Skip to content
Open access

Underwater Computer Vision for Ecologists: A Framework for Curating Image Training Datasets for Object Detection

Sep 2026 · Italian National Conference on Sensors · Vol 26 · 0 citations · 46 references
Medicine

TL;DR

The results suggest that the YOLOv8 object detection model is most effective for abundant and visually distinctive taxa, while performance declines for groups with coarse taxonomic resolution.

Abstract

As marine ecosystems experience accelerating change, there is an urgent need for efficient and scalable biodiversity monitoring tools. We present a 10-step framework for integrating computer vision (CV) tools into long-term underwater biodiversity monitoring, using a case study from coastal British Columbia. Over 9000 h of unbaited remote underwater video footage were collected from two kelp farms and reference sites between March 2022 and June 2023. The framework includes steps for creating an annotated training dataset using unsupervised and supervised CV tools, culminating in the training and validation of a YOLOv8 object detection model. This process produced over 241,000 annotations across 54 pseudo-taxonomic categories (representing both taxa and visually similar groups of fauna), with a focus on fish and gelatinous zooplankton groups. The final model achieved an overall F1 (2 × precision × recall/(precision + recall)) of 0.74 and a mean average precision at 0.5 intersection over union (mAP50) of 0.78. The model had the highest performance on fine-resolution taxa such as Phanerodon vacca (F1 = 0.88) and Aurelia labiata (F1 = 0.90), and lowest performance on pseudo-taxa with limited visual distinctiveness such as Actinopterygii (F1 = 0.60) and Cnidaria (F1 = 0.60). A re-training experiment using annotation thresholds between 25 training images to full dataset availability (~200–24,000 images per group) found that model performance was positively correlated with annotation effort, with F1 averaging 0.84 and mAP50 averaging 0.91 at the maximum training dataset size. Our results suggest that the model is most effective for abundant and visually distinctive taxa, while performance declines for groups with coarse taxonomic resolution. We recommend optimizing annotation effort by targeting genus- or species-level taxa, having at least one broad-level group to capture order-level abundances, and supplementing annotations of rare groups with common but morphologically similar groups, which may further improve model performance.

Read PDF

Similar papers

Open access 2026

A Monocular Depth Estimation Framework for Improved Underwater Fish Biomass Assessment

A computer vision pipeline that integrates monocular depth estimation using MiDaS with YOLOv11 object detection with the aim of delivering accurate and reliable weight estimation without the need for using stereo-cameras is presented, making it very well-suited for practical deployment in real-world aquaculture environ...

Said Al-Abri, S. Keshvari, Rami Al-Hmouz et al. · 0 citations
Review Open access Aug 2026

DeepLitterAI: Automated detection and quantification of deep-sea benthic plastic and macrolitter with field validation in waters around Japan.

Deep-sea imagery is a non-destructive tool for monitoring seafloor litter, but manual inspection limits its scalability. DeepLitterAI, a YOLOv11x-based detector combined with BoT-SORT tracking, was developed to automatically detect and quantify deep-sea benthic macrolitter. The model was trained with J-Litter, a datase...

Ryota Nakajima, Takaki Nishio, H. Saito et al. · 0 citations
Open access Sep 2026

UCOD: A Near-Field Benthic Organism Dataset for Underwater Visual Camouflage and Multi-Task Analysis

Accurate near-field underwater visual perception plays a crucial role in marine ecological monitoring and benthic resource exploration. However, the superimposition of the biomimetic characteristics of benthic organisms and the optical degradation caused by the water medium frequently induces significant underwater v...

Rui-Xue Wang, Xuan-He Chu, Xin-Yu Zhao et al. · 0 citations
#small language model Preprint Sep 2026

Automated Species Identification in Camera Trap Images for Wildlife Conservation

A novel end-to-end framework integrating a self-attention mechanism to address limitations in effectively detecting small animals in low-contrast trap images and small animals while also demonstrating zero-shot detection capability leveraging the MLLM.

Nowshin Amin, Nafisa Tabassum Oyshi, Tahmid Abrar Zidan et al. · 0 citations
Review Open access Aug 2026

WIO-ReefFish: A High-Resolution Dataset for Taxon-Aware Coral Reef Fish Detection in the Western Indian Ocean

WIO-ReefFish, a reef fish detection dataset derived from diver-operated line-intercept transects and designed for ecological monitoring under natural survey conditions, is presented and established as a realistic benchmark for automated reef fish detection and a foundation for more robust computer-vision tools in coral...

J. Gerard, Luca Branger, F. Huyghe et al. · 0 citations
Book Open access Aug 2026

OpenAqua: A Large-Scale Fine-Grained Dataset and Benchmark for Open Underwater Visual Perception

Monitoring aquatic biodiversity is vital for maintaining global ecological balance. While advancements in computer vision have revolutionized underwater perception, existing datasets are predominantly limited to coarse-grained categories or lack spatial localization annotations, severely constraining the applicability...

Linxuan Luo, Pan Mu, Cong Bai · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.