Skip to content

An efficient and robust object detector for complex autonomous driving scenarios

Aug 2026 · Multimedia Systems · Vol 32 · 0 citations · 32 references

TL;DR

Comparisons with mainstream state-of-the-art (SOTA) algorithms confirm that the proposed SDM-RTDETR method strikes a competitive trade-off between detection accuracy and model inference efficiency, demonstrating substantial engineering applicability in complex autonomous driving environments.

View source

Similar papers

#artificial intelligence Review Aug 2026

One-Stage Object Detectors in Autonomous Driving

Autonomous vehicles depend on fast and reliable perception systems to detect surrounding vehicles, pedestrians, cyclists, traffic signs, and other road objects in real time. This paper presents a comprehensive survey and analysis of one-stage object detectors for autonomous driving rather than an implementation of a new detection system. The survey reviews the evolution of major one-stage detectors, including YOLOv1, SSD, RetinaNet, EfficientDet, anchor-free detectors such as FCOS and CenterNet, and recent real-time models such as YOLOv10. It compares these architectures through their design choices, feature-fusion strategies, loss functions, deployment trade-offs, and reported benchmark performance. The paper also summarizes commonly used autonomous-driving datasets, evaluation metrics, open challenges, and future research directions. Overall, this survey highlights how one-stage detectors balance speed, accuracy, efficiency, and robustness, while also emphasizing the remaining gap between benchmark results and dependable real-world autonomous-driving performance.

Jonel Roman, Ryan Sirjue, Peter Nguyen et al. · 0 citations
Open access Aug 2026

EPLS‐YOLO: A Multi‐Scale Object Detection Method for Complex Traffic Scenarios

Aiming at the problems of high missed detection rate and unbalanced feature extraction caused by the coexistence of multi‐scale objects in complex traffic scenarios, this paper proposes an improved YOLO11 detection algorithm (EPLS‐YOLO) for autonomous driving. First, the KITTI dataset is reconstructed and expanded, which is uniformly categorised into six classes (Car, Van, Truck, Pedestrian, Cyclist, and Tram). Data augmentation strategies such as Mosaic and random transformations are further adopted to enhance sample diversity. Second, in the backbone network, the Efficient Multi‐Scale Attention (EMA) is embedded into the C3K2 module to construct the C3K2_EMA module, which strengthens the fine‐grained feature representation of small objects and occluded objects. A P2 high‐resolution detection branch is added to cover the scale of extremely small objects, and a Lightweight Shared Convolution Detection Head (LSCD) is designed to achieve efficient fusion of multi‐scale features from shallow and deep layers. Finally, the SlideLoss function is introduced to dynamically assign sample learning weights based on Intersection over Union (IoU), alleviating the problem of unbalanced training of multi‐scale samples. Experimental results show that the proposed EPLS‐YOLO achieves a precision (P) of 94.5%, recall (R) of 92.0%, and mean Average Precision at IoU = 0.5 (mAP50) of 95.7% on the KITTI dataset, which are 0.7, 3.0 and 1.6 percentage‐points higher than those of the original YOLO11, respectively. Notably, the detection performance for small objects is significantly improved. Moreover, the overall detection performance of EPLS‐YOLO outperforms that of mainstream object detection models such as RT‐DETR, YOLOv8, YOLOv10 and YOLOv12, which can meet the demand for accurate perception of full‐scale objects in autonomous driving.

Tao Feng, Siyi Yang, Wenli Wang et al. · 0 citations
Conference Jul 2026

Robust multitask learning framework for panoptic autonomous driving perception

This study puts forward a resilient cross-task modeling architecture to overcome the precision-versus-latency dilemma in comprehensive traffic scene understanding. The backbone employs a C2f module to enhance gradient flow and small-object detection, while structural re-parameterization via repVGG blocks enables multi-branch training and fast single-path inference. In the segmentation decoder, the CARAFE content-aware upsampling operator replaces nearest-neighbor interpolation to preserve fine edge details, significantly improving lane line continuity. When assessed on the BDD100K corpus, our model's empirical results indicate it surpasses YOLOP as well as HybridNets in detecting objects, segmenting drivable zones, and delineating lane markings. Removal experiments verify the synergistic contributions of each module. Operating at 186 frames per second and utilizing 16.89 million parameters, the network achieves an advantageous tradeoff between precision and computational efficiency for perception tasks in autonomous vehicles.

Qian Luo, Jiangpeng Du, Yawei Li · 0 citations
Open access Jul 2026

LCA-Net: A Lightweight Network for Small Object Detection in Road Traffic Scenes

LCA-Net is presented, a computationally efficient framework for small object detection that balances accuracy with model complexity that demonstrates a favorable accuracy–efficiency trade-off for real-time traffic perception.

Shan Lin, Ben-Sheng Yun, Zhenyu Lin et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.