AHD-YOLO: An efficient front-end perception framework for vision-based drowning-risk monitoring in complex water-surface environments
Abstract
Reliable localization of partially visible humans is an essential front-end requirement for vision-based drowning-risk monitoring. However, horizontal-view water-surface surveillance remains challenging because visible human regions are often small and incomplete, while waves, foam, reflections, motion blur, and lens contamination introduce strong background interference. This paper presents AHD-YOLO, a lightweight detector based on YOLO11n for horizontal-view water-surface human-part detection. The proposed framework does not perform single-frame drowning-event recognition; instead, it detects exposed heads and hands to provide spatial observations for subsequent tracking, temporal behavior analysis, and alarm decision modules. A layer-wise hybrid downsampling strategy combines robust feature downsampling (DRFD) and Haar wavelet downsampling (HWD) to preserve salient semantic responses and high-frequency boundary details at different feature levels. In addition, an asymmetric cross-domain attention (ACA) module recalibrates multi-scale fused features to enhance target-related responses and suppress water-surface interference. A dedicated dataset containing 1,222 source images is reannotated with head and hand bounding boxes and expanded to 6,110 images using reflection, water stain, camera shake, and mixed degradations. Experimental results show that AHD-YOLO achieves 92.47% mAP50\documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{mathrsfs} \usepackage{upgreek} \setlength{\oddsidemargin}{-69pt} \begin{document}$$\hbox {mAP}_{50}$$\end{document} and 57.18% mAP50--95\documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{mathrsfs} \usepackage{upgreek} \setlength{\oddsidemargin}{-69pt} \begin{document}$$\hbox {mAP}_{50\text {--}95}$$\end{document} with 3.06M parameters and 3.5 GFLOPs. These results demonstrate that the proposed model provides an effective accuracy–complexity trade-off for front-end perception in real-time drowning-risk monitoring.