Jul 2026· Engineering Research Express· Vol 8, pp. 145228· 0 citations· 43 references
Physics
TL;DR
AD-YOLO is presented, a dual-level framework that tackles pseudo-label noise and limited multi-scale adaptability when applied to semi-supervised object detection frameworks from both the detector architecture and the SSOD pipeline.
Abstract
In intelligent transportation systems, real-time detection demands high inference speed, making single-stage detectors the preferred choice. However, existing semi-supervised object detection (SSOD) frameworks suffer from pseudo-label noise and limited multi-scale adaptability when applied to such detectors. This paper presents AD-YOLO, a dual-level framework that tackles these issues from both the detector architecture and the SSOD pipeline. At the detector level, a normalization-guided attention module enables feature recalibration with zero extra parameters, and a CIoU-NWD hybrid loss incorporates the Wasserstein distance to suppress localization jitter. At the framework level, an Adaptive Teacher employs a category-aware dynamic threshold that selects pseudo-labels based on confidence percentiles, eliminating preset thresholds and calibration bias. This thresholding is assisted by a scale-aware dynamic augmentation mechanism, which uses teacher-generated pseudo-labels to identify small-object images and applies only weak augmentation to preserve their semantics. These two levels form a self-reinforcing loop. Experiments on BDD100K and TT100K show that with only 10% labeled data, AD-YOLO achieves 28.83% mAP at 142 FPS, demonstrating better performance compared with peer methods and exhibiting robustness in complex traffic scenes.
An Adaptive and Scalable YOLO model named AS-YOLOR (Adaptive and Scalable YOLO for Rotated object detection), based on the YOLOv8 baseline is proposed, providing a solution with strong practical potential for achieving efficient and high-precision detection of small, rotated objects.
Jin Huang, Juntao Shen, Min Wang et al.· Applied Sciences· 0 citations
In the label noise suppression strategy (LNSS), a contrastive learning mechanism based on semantic consistency is introduced to constrain the aggregation of similar samples in the feature space, thereby reducing the adverse impact of noisy samples on model optimization.
Nuo Chen, Peng Zhao, Shouquan Hou· Remote Sensing· 0 citations
The proposed framework features a redesigned cross-scale feature fusion module, CCFM-S2, which utilizes the SPD-Conv operator for information preserving downsampling and explicitly integrates high-resolution shallow features (S2 layer), thereby infusing indispensable spatial details into the feature hierarchy for small targets.
Real-time object detection requires identifying objects in video streams or consecutive images with minimal latency, yet it continues to struggle with non-salient objects—those that are small, occluded, or otherwise inconspicuous. To address this limitation, this paper proposes CAM-YOLO, an enhanced architecture based on YOLOv8. First, to mitigate the baseline model’s limited representational capacity for non-salient targets, we introduce a Multi-Scale Aggregation Module (MSAM) into the feature fusion process, enabling the backbone network to extract more discriminative fine-grained features. Second, to better capture global contextual relationships associated with such objects, we design a Contextual Association Module (CAM) that explicitly models long-range spatial dependencies. Furthermore, we integrate a Dual-Branch Attention Mechanism (DBAM) to refine the feature processing flow, thereby strengthening the contextual feature representations crucial for detecting non-salient instances. Extensive experiments on two large-scale public benchmarks, Microsoft Common Objects in Context 2017 (MS COCO 2017) and PASCAL Visual Object Classes (PASCAL VOC), demonstrate that CAM-YOLO achieves highly competitive performance compared to several widely-adopted realtime detectors.OPEN ACCESS Received: 19/11/2025 Accepted: 15/01/2026 Published: 21/07/2026
C. Dong, Y. Ding, J. Hu· Revista Internacional de Mét...· 0 citations
The rapid advancement of deep learning has enabled intelligent analysis in professional sports, yet tennis remains particularly challenging due to small and fast-moving objects, frequent occlusions, and complex backgrounds. To address these difficulties, we propose YOLO-Net, a lightweight detection framework tailored for tennis event analysis. Built upon YOLO11n, the framework integrates three task-oriented improvements: a C3k-MSEIS module for multi-scale edge enhancement and dual-domain feature selection to refine fine-grained boundaries; an ECA channel attention mechanism inserted after C2PSA to strengthen inter-channel dependency modeling and improve feature discriminability; and a Focaler-IoU loss function to emphasize hard and small samples while reducing localization errors. In addition, we construct and annotate a dedicated tennis dataset containing 6,648 images across three categories—player, racquet, and ball—covering diverse scenes, camera angles, and lighting conditions. Experimental results show that YOLO-Net achieves 84.5% precision and 78.2% mAP@0.5 with only 2.58M parameters, outperforming the YOLO11n baseline by 2.5% in precision and 0.9% in mAP while maintaining real-time inference. These findings demonstrate that YOLO-Net is an efficient, accurate, and deployable solution for applications such as referee assistance, tactical analysis, and intelligent broadcasting in tennis competitions.
Xiangyu Du, Tao Wang, Weiwei Zu et al.· PLoS ONE· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.