Similar papers
A multimodule collaborative optimization method for YOLOv11-based vehicle detection in complex traffic scenes
To address the problems of weak small target features, target occlusion, severe weather interference, and bounding box regression bias in vehicle detection under complex traffic scenarios, a multi-scale feature fusion vehicle detection method is proposed. An improved YOLOv11 model is used for vehicle target detection. First, data augmentation techniques are used to expand the sample size for complex weather and lighting distortion scenarios to improve the model's environmental adaptability. Second, an efficient multi-scale attention mechanism (EMA) is embedded in the backbone network to enhance target feature extraction and suppress background redundancy. Third, a neck network using PSConv replaces the traditional downsampling convolution to optimize small target feature preservation and multiscale receptive field configuration (MSRF). Finally, an Inner-EIoU loss function is introduced during the training phase to improve bounding box regression accuracy and training stability. Experimental results show that on the UA-DETRAC benchmark dataset and the self-built complex traffic scene dataset, the mAP50 and mAP50-90 values of the vehicle detection model in complex scenes are 64.8% and 46.5%, respectively, which are improved by 4.5% and 4.2% by YOLOv11. From the perspective of computational efficiency, the improved model has a GFLOPs of 6.60, which is lower than YOLOv11's 6.90, and has better detection performance.
SRIN-YOLO: Algorithms for High-Altitude Drone Detection
In response to the problems of easy missed detection, false detection, and limited detection performance in multi-scale and small target scenes of UAVs in complex backgrounds at high altitudes, this paper proposes a lightweight target detection model SRIN-YOLO based on YOLOv11n. In order to enhance the multi-scale feature modeling ability in complex backgrounds, this paper designs a lightweight attention mechanism iREMA that combines inverted residual structure and EMA module to enhance the correlation modeling between features of different scales, thereby improving the feature expression of small targets. In addition, a dynamic weighting strategy is introduced under the Shape-NWD loss function framework, so that the model can adaptively adjust the training weight according to the sample difficulty, so as to pay more attention to the difficult samples and small targets in the training process, and improve the convergence efficiency and detection accuracy of the model. Finally, in terms of network structure, a lightweight backbone and a small target detection head are adopted to achieve a balance between detection accuracy and inference efficiency. The experimental results show that the parameter quantity of SRIN-YOLO is only 1.24 M (2.69MB), which is 51.69% less than that of YOLOv11n. On the TIB-Net dataset, the model achieved 93.2%, 92.5% and 92% on the Precision, Recall and mAP@0.5 indicators, respectively, which were 3%, 13% and 5% higher than the baseline model, respectively, and the inference speed reached 188 FPS. To verify the generalization ability of the model, this paper further conducts experiments on the VisDrone dataset. The results show that SRIN-YOLO is superior to the baseline model in multiple evaluation indicators. The above experimental results show that the proposed method achieves a good balance between detection accuracy and computational efficiency, and is suitable for resource-constrained embedded platforms and real-time drone vision application scenarios.
YOLOv8 vehicle detection method based on lightweight coordinate attention
To address the challenges encountered in vehicle detection within complex traffic scenarios, such as diverse features, strong background interference, and deployment constraints on edge devices, a lightweight vehicle detection method based on YOLOv8 improved with the Coordinate Attention mechanism is proposed. In this method, the Coordinate Attention module is introduced into the YOLOv8s backbone network to reconstruct the C2f structure. By embedding positional information, the model enhances its capability to extract key features, suppressing complex background interference while maintaining its lightweight characteristics. Experiments conducted on a dataset comprising 17,428 images of 17 vehicle classes demonstrate that the improved model achieves a precision of 90.4% and an mAP@0.5 of 91.5%. For vehicle types with regular structures and large volumes, such as four-wheel small trucks and six-wheel medium trucks, the detection accuracy approaches 99%, and the detection confidence remains stable above 0.8 under complex lighting and background interference conditions. This method effectively controls model complexity while maintaining high detection accuracy, thereby satisfying the application requirements for real-time vehicle detection on edge devices such as intelligent driving recorders.
SFC-YOLO: An Accuracy-Enhanced and Parameter-Efficient YOLO Framework for Small Vehicle Detection in Aerial Images
Small vehicle detection in aerial images is important for intelligent transportation, low-altitude inspection, urban monitoring, and vision-based sensing systems. Vehicles in aerial images often occupy few pixels and are affected by complex backgrounds, shadows, viewpoint changes, weak texture, and similar class appearances, which can cause missed detections and false alarms. To address these issues, this paper proposes SFC-YOLO, an accuracy-enhanced and parameter-efficient framework based on YOLO11n. A Feature Complementary Block (FCB) is placed at the P5/32 high-level feature stage to enhance local-detail and semantic-context compensation; a Dynamic Feature Alignment Upsampling unit (DFAU) is inserted into the first P5-to-P4 upsampling path to improve content-adaptive feature alignment; and a P5-only Hidden-State Attention (HSA) module is used in the final P5 detection branch to strengthen global semantic interaction. Experiments on the VEDAI eight-class vehicle dataset show that SFC-YOLO reduces the parameter count from 2.584 M to 2.466 M while improving the five-run mean Precision from 0.602 to 0.665, Recall from 0.599 to 0.608, and mAP50 from 0.614 to 0.652. Since the GFLOPs increase from 6.3 to 8.4, SFC-YOLO should be interpreted as a parameter-reduced but not FLOP-reduced framework. The main trade-off is improved detection accuracy with fewer trainable parameters at the cost of higher theoretical computation. Five repeated experiments yield an average mAP50 of 0.6523 ± 0.0117, quantifying the run-to-run variation under the current training protocol.
An enhanced DEIM-based method for target detection in complex field environments
To address the issues of complex background interference, similar target misdetection and low recognition accuracy of tiny objects in Unmanned Aerial Vehicle (UAV) object detection scenarios, this study proposes an enhanced DEIM-based method for target detection in complex field environments. First, the backbone network is integrated with an Omni-Dimensional Dynamic Convolution (ODConv) module, which alleviates false detection caused by the indistinct feature differences between similar distractors and objects by exploiting the multi-dimensional dynamic attention learning mechanism in the convolution kernel space. Second, a weighted convolution-based C3k2 (wConv-C3k2) module is constructed in the encoder network, which suppresses redundant background information and reduces missed detection caused by background noise by introducing a spatial density function to dynamically adjust the weights of the convolution kernel. Finally, a Deep Robust Feature Downsampling (DRFD) module is designed, which reduces the loss of key features and enhances the capability of small object detection by fusing three branches with complementary characteristics. Experimental results demonstrate that compared with the baseline DEIM model, the proposed model yields 3.5%, 4.9%, and 3.0% improvements in AP50:95, AP75, and AP50, respectively, together with a 5.4% increase in mAR. This method effectively improves the accuracy and reliability of target detection in aerial images, enriches the technical approaches of intelligent control and information perception, and provides reliable technical support for the practical applications of information and control systems in field rescue.
YOLOv8n-MMSP: A Lightweight Real-Time Vehicle Detector for Edge-Based Perception in Intelligent Transportation Systems
Real-time vehicle detection in edge-based intelligent transportation systems is essential for enhancing traffic efficiency and safety. However, vehicle detection in complex urban environments faces challenges (cluttered backgrounds, significant scale variations among targets), necessitating lightweight models that achieve both high accuracy and low computational cost for resource-constrained edge devices. To address these challenges and tackle key engineering problems in traffic scenarios, we proposed YOLOv8n-MMSP, a lightweight real-time vehicle detection algorithm optimized for edge deployment. To mitigate fine-grained feature loss from downsampling, the C2f-StarBlocks module was integrated to enhance multiscale feature fusion, beneficial for long-range vehicle perception in traffic monitoring systems. To reduce the distant small-object miss rate, a P2 small-object detection layer was introduced to strengthen shallow feature representation, improving detection reliability. To handle significant scale variations and complex background interference, the Multi-Scale Convolutional Attention mechanism was incorporated, enabling robust perception under challenging conditions (nighttime and adverse weather). To improve bounding box regression accuracy in densely occluded scenarios, the Multi-Precision Distance Intersection over Union loss function was adopted, facilitating more precise localization in congested traffic environments. Furthermore, to meet the computational constraints of edge devices, an L1 regularization–based channel pruning strategy was applied to batch normalization layers, reducing model complexity while maintaining detection performance. Results demonstrated 81.2% mAP and 61.4% mAP50-95, consistent improvements over the baseline model, while reducing model size, parameter count, and computational load to 61.89%, 59.12%, and 58.40% of the original, respectively. Edge-device tests further demonstrate a real-time inference speed of 73.6 fps.