A lightweight YOLO-style detector that integrates a PP-HGNetV2 tiny backbone, an enhanced normalization-based attention module (ImNAM), and an improved complete intersection-over-union loss (ImCIoU) shows strong potential for intelligent industrial inspection on resource-constrained platforms, subject to further hardware-level deployment verification.
Abstract
Mechanical component detection in industrial scenes is challenged by cluttered backgrounds, large-scale variation, specular reflection, high inter-class similarity, and class imbalance. To address the above problems, this paper proposes a lightweight YOLO-style detector that integrates a PP-HGNetV2 tiny backbone, an enhanced normalization-based attention module (ImNAM), and an improved complete intersection-over-union loss (ImCIoU). The HGNetV2 backbone enhances hierarchical multi-scale feature extraction and keeps the deployable computational complexity low. ImNAM has been modified to enhance discriminative representation by introducing dual-statistics channel weighting, orthogonal edge-aware spatial modeling and bipolar adaptive residual gating. ImCIoU enhances the accuracy of localization by combining quality-aware box scaling, scale-sensitive modulation and dynamic IoU-guided weighting. A class-balancing augmentation pipeline was applied to the four-category industrial dataset of Bearing, Bolt, Gear and Nut. All experimental results are reported as the mean ± standard deviation of five independent two-tailed training runs with different random seeds, and statistical significance is verified by paired t-tests (p < 0.05) with Bonferroni correction for multiple comparisons. Experimental results show that the proposed method achieves 90.82 ± 0.35% mean average precision (mAP@0.5), 91.95 ± 0.42% precision, and 82.98 ± 0.51% recall, outperforming nine mainstream lightweight detectors, including the latest YOLOv12n (2025) and RT-DETR-tiny. Extended evaluation on mAP@0.5:0.95, per-class AP and F1 score further confirms the advantages in localization accuracy and classification performance. Ablation studies confirm that the HGNetV2 family backbone provides the largest performance gain, while the improved attention mechanism and regression loss further enhance localization accuracy and robustness. With only 4.44 M parameters and 9.96GFLOPs, the proposed detector has achieved a good accuracy–efficiency trade-off and shows strong potential for intelligent industrial inspection on resource-constrained platforms, subject to further hardware-level deployment verification.
A lightweight object detection framework, termed MDCF-YOLO, which achieves a superior accuracy-efficiency trade-off compared to state-of-the-art lightweight object detectors and exhibits similarly competitive performance on the AI-TOD dataset, further validating its effectiveness and generalization capability in UAV remote sensing scenarios.
Peng-Fei Dai, Liang Chen, Ting Fan et al.· Cluster Computing· 1 citation
Surface defects generated during steel manufacturing significantly affect product quality, structural reliability, and operational safety, creating a strong demand for accurate and real-time inspection systems in industrial environments. However, existing detection approaches often struggle with subtle defect textures, complex surface backgrounds, weak visual contrast, and the trade-off between detection accuracy and computational efficiency. To address these challenges, this paper proposes SDM-YOLO, a lightweight framework for steel surface defect detection based on YOLO11n. Rather than introducing an entirely new detection architecture, the main contribution of this work lies in the coordinated integration and task-specific adaptation of complementary modules within a unified lightweight detection framework. Specifically, the proposed method enhances feature representation by replacing the original C2PSA module with C2PSA_SEAM in the backbone, introduces DySample-based dynamic upsampling in the neck for content-aware multi-scale feature alignment, and incorporates a multi-scale convolutional attention mechanism before the detection head to improve sensitivity to subtle, low-contrast, and morphologically varied defects. In addition, a normalized Wasserstein distance loss is employed to improve localization stability for small and overlapping defects without increasing inference-time parameters or computational cost. Extensive experiments on the NEU-DET and GC10-DET datasets demonstrate that SDM-YOLO achieves mAP50 scores of 81.0% and 72.3%, respectively, while attaining mAP50:95 values of 46.7% and 38.0%. The proposed framework maintains real-time performance with only 2.68 M parameters, 6.6 GFLOPs, and an inference speed of 94.5 FPS. These results demonstrate that SDM-YOLO achieves an effective balance between detection accuracy, localization precision, and computational efficiency, making it suitable for practical steel surface defect inspection applications.
Nabin Kandel, Ping Wu· Engineering Research Express· 0 citations
The rapid advancement of deep learning has enabled intelligent analysis in professional sports, yet tennis remains particularly challenging due to small and fast-moving objects, frequent occlusions, and complex backgrounds. To address these difficulties, we propose YOLO-Net, a lightweight detection framework tailored for tennis event analysis. Built upon YOLO11n, the framework integrates three task-oriented improvements: a C3k-MSEIS module for multi-scale edge enhancement and dual-domain feature selection to refine fine-grained boundaries; an ECA channel attention mechanism inserted after C2PSA to strengthen inter-channel dependency modeling and improve feature discriminability; and a Focaler-IoU loss function to emphasize hard and small samples while reducing localization errors. In addition, we construct and annotate a dedicated tennis dataset containing 6,648 images across three categories—player, racquet, and ball—covering diverse scenes, camera angles, and lighting conditions. Experimental results show that YOLO-Net achieves 84.5% precision and 78.2% mAP@0.5 with only 2.58M parameters, outperforming the YOLO11n baseline by 2.5% in precision and 0.9% in mAP while maintaining real-time inference. These findings demonstrate that YOLO-Net is an efficient, accurate, and deployable solution for applications such as referee assistance, tactical analysis, and intelligent broadcasting in tennis competitions.
Xiangyu Du, Tao Wang, Weiwei Zu et al.· PLoS ONE· 0 citations
The identification of small targets in aerial imagery continues to present difficulties owing to variations in object scale and insufficient feature characterization. Current methodologies enhance detection accuracy through attention-based mechanisms or multi-level feature integration, yet frequently result in substantial computational demands. This study introduces QDP-YOLOv5, an enhanced detection architecture derived from HIC-YOLOv5, specifically designed for effective small target recognition. The proposed framework incorporates an innovative Query-Guided Deformable Pyramid (QDP) component that dynamically adjusts perceptual ranges by combining content-sensitive queries with adaptable convolution operations. Distinct from conventional deformable convolution methods, the QDP module leverages semantic queries to guide geometric transformation and spatial sampling, rather than relying solely on local feature-driven offset prediction. These QDP components are strategically embedded across multiple feature hierarchies (backbone endpoint, P4 layer, P3 layer) to strengthen multi-scale feature extraction. Furthermore, the system employs Efficient Channel Attention (ECA) to optimize channel-specific feature weighting with negligible computational burden. Comprehensive ablation experiments validate the optimality of the QDP module’s deployment strategy in the feature pyramid. Evaluation on the VisDrone2019 benchmark reveals that QDP-YOLOv5 outperforms the original HIC-YOLOv5 model by 1.73% in mAP@0.5 and 1.00% in mAP@ [0.5:0.95] metrics, while maintaining a competitive inference speed of 62.3 FPS on an NVIDIA RTX 4090. confirming the proposed method's superior performance.
Feiyu Zhao, Hongchang Ding, Menghao Yang· Digital Signal and Computer...· 0 citations
Visual object detection is essential for environment perception in intelligent robots, automated assembly, unmanned inspection, and industrial detection systems. Although lightweight detectors reduce complexity through compact architectures, fixed convolutional units, and progressive downsampling, their limited scale responses and spatial-detail loss may degrade the localization of scale-varying and boundary-sensitive objects. To address this issue, this paper proposes DGMS-YOLO, an information-preserving lightweight detector built on a one-stage framework. The model improves feature representation through adaptive scale selection and information preservation. Specifically, the Dynamic Gated Multi-Scale Selection (DGMS) module extracts multi-scale features using depthwise convolution branches with different receptive fields and generates content-aware scale weights from the mean and standard deviation of input features. A temperature-scaled softmax and uniform scale prior are further introduced to prevent premature branch-weight concentration during multi-branch training. The Dual Pooling Downsampling (DPD) module combines max pooling, average pooling, and stride convolution to preserve salient responses, regional structures, and learnable semantic features during downsampling. In addition, high-frequency residual calibration estimates edge residuals from low-frequency features and applies lightweight channel gating to compensate for localization-related textures and boundary details. On PASCAL VOC 2007, DGMS-YOLO achieves $\text{7 6. 6 7 \%} \text{m A P}_{50}$ and $\text{5 5. 3 3 \%} \text{m A P}_{50: 95}$ with a parameter count comparable to YOLOv8s, improving it by 1.23 and 2.00 percentage points, respectively. These results demonstrate that dynamic scale selection and information preservation improve detection accuracy and localization quality under a parameter and storage budget comparable to YOLOv8s.
Xuebing Yue, Meng-Kui Hao, Yao Yao et al.· International Conference on...· 0 citations
Infrared (IR) object detection poses unique challenges due to low texture, weak edges, sensor noise, and a strong dependence on global context. Modern convolutional object detectors, including YOLOv8, are primarily optimized for RGB imagery and often underperform when directly applied to large-scale infrared datasets. In this paper, we present a simple yet effective architectural enhancement to YOLOv8 by (1) replacing the standard Spatial Pyramid Pooling Fast (SPPF) module with a learnable multi-scale variant (SPPFPlus), and (2) integrating the Convolutional Block Attention Module (CBAM) into both the backbone and neck. The proposed approach introduces explicit parallel multi-scale context modeling and adaptive channel–spatial attention, addressing key representational limitations of infrared imagery. Extensive experiments on a large infrared dataset demonstrate an improvement of approximately 15% in detection performance over the YOLOv8 baseline, with notable gains in recall and small-object detection. The modifications are lightweight, modular, and can be seamlessly integrated into existing YOLOv8
A. Amankwah· Signal & Image Processin...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.