Jul 2026· 2026 6th International Conference on Electrical, Computer and Energy Technologies (ICECET)· pp. 1-6· 0 citations· 16 references
Abstract
Real-time weapon detection in video surveillance systems is a critical requirement for proactive security applications, particularly under the computational and latency constraints imposed by edge artificial intelligence deployments. While the YOLO family of object detectors has undergone continuous architectural evolution, the recently introduced YOLOv26 represents a significant redesign aimed at improving efficiency, stability, and deployment suitability across a wide range of hardware platforms. This work presents a comprehensive and homogeneous experimental evaluation of the full YOLOv26 model family, ranging from nano (YOLOv26n) to extra-large (YOLOv26x) variants, for real-time weapon detection in surveillance imagery. All models are trained and evaluated under identical conditions using a dataset that explicitly includes visually similar non-weapon objects as hard negatives, enabling a realistic assessment of false positives and false negatives in safety critical scenarios. The analysis encompasses training and validation dynamics, precision, recall evolution, mean Average Precision (mAP) at multiple IoU thresholds, class-wise confusion matrices, and inference latency. Results show that performance improves consistently from smaller to medium sized models, with YOLOv26m achieving the most balanced trade off between detection accuracy, robustness, and computational cost. Larger variants provide marginal accuracy gains at significantly higher complexity, revealing diminishing returns for edge oriented deployments. Overall, the findings demonstrate that the YOLOv26 architecture offers a scalable and mature detection framework, where model selection can be guided by explicit operational criteria rather than raw accuracy alone. This study establishes a strong baseline for future work on real world edge deployment, multi camera surveillance systems, and hardware aware optimization of next generation YOLO detectors.
This study proposes an Advanced Surveillance Framework that makes use of YOLOv10, a next-generation real-time object detection algorithm that greatly outperforms conventional single-sensor approaches in precision, recall, and real-time responsiveness.
Sadiya Begum, Lubna Nausheen, Ruqiya Fatima· International Journal of Eng...· 0 citations
The paper presents a customized version of the YOLOv12 model that enables better detection of small, occluded, and low-contrast weapons in video sequences while maintaining high precision and real-time inference speed. The new model integrates: 1) loss reweighting strategy that emphasizes small objects’ contributions during training; and 2) set of lightweight, append-only enhancement modules placed in the detection head of the baseline architecture. A large-scale custom dataset has been developed, including more than 26,528 images and 38,167 labeled instances, extracted from about 1,200 YouTube videos and curated web images. The dataset includes three weapon types (knife, pistol, long-gun) and an additional no_weapon class (showing images prone to being easily confused with true instances) to reduce false-positive detections. When benchmarked against the baseline YOLOv12s model, our customized version demonstrated a 4.9% increase in mAP@50, 7.2% in mAP@50-95, 3.8% and 7.1% improvements in Precision and Recall, respectively, while maintaining real-time inference speed.
Constantin Catargiu, I. Ciocoiu· IEEE Access· 0 citations
Recent YOLO-based object detectors provide a strong accuracy–latency trade-off for security-critical applications, yet it remains unclear whether attention mechanisms consistently improve modern architectures. This paper presents a controlled ablation study of integrating the Convolutional Block Attention Module (CBAM) into YOLOv8 and YOLOv9 for multi-class surveillance object detection. Experimental results reveal a clear architecture-dependent behavior. CBAM improves YOLOv8 performance, increasing precision from 0.7786 to 0.8060 and mAP@50 from 0.8571 to 0.8753 with modest computational overhead. In contrast, CBAM significantly degrades YOLOv9 performance, reducing mAP@50 from 0.9501 to 0.7600 and mAP@50–95 to 0.5400, while nearly doubling the model size (25.3–48.7 M parameters). These findings demonstrate that attention mechanisms are not universally beneficial; rather, their effectiveness depends on the underlying architecture. In high-capacity models such as YOLOv9, CBAM introduces redundancy that negatively impacts performance. Additional analysis, including efficiency evaluation and Grad-CAM visualization, further supports this interpretation. Overall, this study highlights that attention should be treated as an architecture-dependent design choice rather than a generic performance enhancement, providing practical insights for the development of efficient real-time detection systems.
Debolina Ghosh, J. Singh· Discover Computing· 0 citations
Real-time weapon detection is a critical component of intelligent surveillance systems, particularly for perimeter monitoring applications on embedded edge platforms. However, reliable alarm generation remains challenging because false positives, temporal instability, and viewpoint inconsistencies can propagate through conventional multi-camera fusion strategies. To address these limitations, this work proposes a lightweight Adaptive Multi-Camera Temporal Fusion (ACTF) framework that combines confidence-aware evidence separation, temporal persistence, and short-window cross-camera validation at the decision level, thereby confirming detections without requiring additional neural-network inference. The framework was evaluated using a TensorRT-optimized YOLO26s detector in controlled dual-camera scenarios involving clear visibility, partial occlusion, visually ambiguous distractors, and challenging illumination. While logical OR fusion achieved higher recall, it also propagated erroneous detections; in contrast, ACTF completely suppressed the distractor-induced false alarms while maintaining competitive performance and sub-second confirmation under favorable conditions. The original NVIDIA Jetson Nano implementation achieved an average throughput of 2.4 camera-pair cycles per second, corresponding to low-rate online embedded operation, whereas an additional NVIDIA Jetson Xavier NX benchmark achieved an average of 10.2 camera-pair cycles per second. This is equivalent to 10.2 processed frames per second for each camera stream and 20.4 camera images per second in aggregate. These results support reactive real-time embedded operation on the Xavier NX platform and demonstrate that ACTF improves alarm reliability with negligible decision-level computational overhead.
Carlos Julio Fierro-Silva, Carolina Del-Valle-Soto, S. M. Mostafa et al.· IEEE Access· 0 citations
This project proposes an AI-powered threat detection system capable of automatically detecting weapons in real time from CCTV footage, specifically focusing on pistols. Security in modern society is a growing concern, especially for countries aiming to create a safe environment for investors and tourists. While Closed Circuit Television (CCTV) cameras are widely used for surveillance, they still depend on human oversight. This project addresses the need for an automated system capable of detecting illegal activities, specifically weapons, in real-time using CCTV footage. Current deep learning techniques, despite advancements in hardware and software, face challenges such as occlusions, viewing angles, and varied environments. This project proposes a solution utilizing state-of-the-art deep learning algorithms for weapon detection, focusing on pistol detection using a custom dataset created from various sources, including manual collections, YouTube, GitHub, and public databases. Through testing multiple models, YOLOv4 emerged as the most effective with an F1-score of 91% and mean average precision (mAP) of 91.73%.
Keywords--Weapon Detection, Deep Learning, Object Detection, Artificial Intelligence, Computer Vision, Real-time Surveillance
Dr.Pothuraju V V Satyanarayana, Chittiboina Syamala· International Scientific Jou...· 0 citations
Tests show that the proposed SmartVision-AI architecture can deliver face recognition accuracy, multi-class weapon detection accuracy, and a precision increase of up to 19%, and a processing time of only 38 ms/frame, which can be effectively deployed in near real-time.
P.Shobana, V. S. Raja, P. S. Rajakumar et al.· International journal of com...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.