Skip to content
Preprint

City Sentinel: A Unified AI-Based Smart Surveillance Framework for Real-Time Multi-Threat Detection Using Deep Learning

Aug 2026 · 0 citations · 15 references
Computer Science

TL;DR

The results demonstrate that a modular, open-source, multi-model architecture can provide broad surveillance coverage, cloud-based auditability, and flexibility for adding new detection capabilities while maintaining practical real-time performance.

Abstract

Rapid urbanization has increased the need for surveillance systems that can monitor multiple public safety risks at the same time. Traditional systems often use separate solutions for facial recognition, vehicle identification, fire detection, and behavioral analysis, resulting in fragmented infrastructure and multiple interfaces for operators to manage. This paper presents City Sentinel, a unified AI-based surveillance framework that integrates six detection capabilities into one scalable platform: facial recognition, automatic number plate recognition (ANPR), fire and smoke detection, weapon and knife detection, violence detection, and road accident detection. The system combines a Next.js operator dashboard, FastAPI backend, cloud-based PostgreSQL event storage, InsightFace and YOLOv8 vision models, and EasyOCR for plate recognition. Camera streams are processed through dedicated inference workers using RTSP. On a workstation equipped with an NVIDIA RTX 3060 GPU, the system achieves a median end-to-end latency of 743 ms and supports four concurrent RTSP streams within a two-second latency limit. It achieves a 91.2% face-match rate, 85.7% plate-reading accuracy, and mAP@0.5 scores of 0.846 to 0.889 across the fire, knife, and weapon detection modules. In user-acceptance testing, operators could enroll a new identity in under one minute and identify a flagged person from live footage in an average of 12 seconds. The results demonstrate that a modular, open-source, multi-model architecture can provide broad surveillance coverage, cloud-based auditability, and flexibility for adding new detection capabilities while maintaining practical real-time performance.

View source

Similar papers

Open access Aug 2026

REAL-TIME WEAPON AND THREAT DETECTION USING YOLOV12 WITH MULTI-SENSOR FUSION FOR ENHANCED SURVEILLANCE SYSTEMS

This study proposes an Advanced Surveillance Framework that makes use of YOLOv10, a next-generation real-time object detection algorithm that greatly outperforms conventional single-sensor approaches in precision, recall, and real-time responsiveness.

Sadiya Begum, Lubna Nausheen, Ruqiya Fatima · 0 citations
Open access Jul 2026

An Intelligent Deep Learning Framework for Real-Time Weapon Detection in Smart Surveillance Systems

DeepGuard is presented, an intelligent deep learning framework for automated weapon detection in images and surveillance videos using Faster Region-Based Convolutional Neural Network (Faster R-CNN) and Single Shot Detector (SSD).

Chengoli prashanth, Sk.Mahammadunnisa · 0 citations
Conference Open access 2025

Smart Campus Surveillance System: A Multi-Modal AI Approach to Real-Time Threat Detection and Response

The proposed framework unifies the fight and weapon recognition, face identification and contextual interpretation of events into a unified monitoring pipeline, and a novel contribution of this work is the tool calling that allows the VLM to automatically seize important frames and trigger alert protocols in such a way that it will reduce man-in, and response delays.

M. Kurulekar, Sanjesh Pawale, Tanay Ingale et al. · 0 citations
Open access Jul 2026

SMARTVISION-AI: A UNIFIED DEEP LEARNING ARCHITECTURE FOR FACE RECOGNITION AND WEAPON DETECTION IN CCTV VIDEO STREAMS

Tests show that the proposed SmartVision-AI architecture can deliver face recognition accuracy, multi-class weapon detection accuracy, and a precision increase of up to 19%, and a processing time of only 38 ms/frame, which can be effectively deployed in near real-time.

P.Shobana, V. S. Raja, P. S. Rajakumar et al. · 0 citations
Open access Aug 2026

SmartFire Vision: An Attention-Pruned Hybrid Vision Transformer and Detection Transformer Framework for Accurate, Efficient, and Real-Time Fire and Smoke Detection in Smart City Video Surveillance

Fire incidents can lead to significant destruction of lives and property, especially in urban and smart cities, and pose a great risk worldwide. Existing fire and smoke detection systems are often inadequate for detecting the location of a fire, assessing the speed of its spread, and providing real-time alerts that can be acted upon quickly. This study proposes a method termed SmartFire Vision, which uses a hybrid deep learning framework consisting of an Efficient Vision Transformer (E-ViT) and a Detection Transformer (DETR) for real-time fire and smoke detection from video sequences. A major contribution of this study is the integration of a new Removing Inefficient Attention Heads (RIAH) pruning strategy to reduce the computational overhead and maintain a global context in the ViT encoder. The E-ViT and DETR feature representations were fused and passed to a fully connected classification head enhanced with a probabilistic thresholding function and an integrated alarm system. The proposed model was trained and evaluated using the FURG fire benchmark dataset, which comprises 28,022 annotated frames. The proposed model achieved an overall accuracy of 85.40%, precision of 85.33%, recall of 85.43%, and F1-score of 85.35%, surpassing the current state-of-the-art methods. The SmartFire Vision framework provides a highly capable and computationally efficient means of fire detection, is particularly beneficial for CCTV-based smart city surveillance, and shows promising computational efficiency on desktop-class GPUs, though dedicated edge-hardware validation remains a direction for future work.

Muhammad Azhar, Muhammad Arman, Asma Iqbal et al. · 0 citations
Conference Jul 2026

YOLOv8-Powered Intelligent Surveillance: An Integrated Real-Time Framework for Crowd Management, Crime Prevention, and Workplace Safety Monitoring using AI and ML

The evolving complexity of urban environments and the effectiveness of traditional CCTV surveillance is making it increasingly difficult to ensure public safety, effective crowd management, crime prevention, and workplace security solutions. However, the traditional approach to surveillance is largely manual, leading to late reactions, missed events, and scalability issues. The intelligent surveillance system based on the YOLOv8 object detection algorithm is designed to enhance workplace safety, prevent crimes, manage crowds, and achieve face recognition in real time within a single AI platform. This paper introduces the concept of an intelligent surveillance framework that combines real-time face recognition, workplace safety monitoring, crowd management, and crime prevention through the use of YOLOv8 object detection algorithms within a single AI-driven solution. The framework employs the YOLOv8 algorithm for object detection, identifying people, weapons, suspicious activities, abandoned objects, and workplace safety violations, and issuing automatic alerts to facilitate swift decision-making. The proposed framework provides an integrated platform of multiple surveillance functionalities as opposed to the existing surveillance systems, where each surveillance task is monitored separately, which provides overall situational awareness using the existing CCTV. The model was trained with surveillance images annotated and tested with Precision, Recall, F1-score, Accuracy, and mAP@0.5. An overall detection accuracy of 92.4%, a precision of 92.4%, a recall of 89.7%, an F1-score of 91.0%, and an mAP@0.5 of 93.2% have been achieved during experimental evaluation. Moreover, the framework's average inference latency is 18ms per frame, which guarantees that it can be used in real-time surveillance applications without compromising the accuracy of its detection results when deployed in various surveillance environments. The proposed system is versatile and feasible for implementation in smart city systems, transportation hubs, industrial production sites, and various organizational environments, and can enable smart surveillance by merging multiple security functions into a single framework based on the YOLOv8 object detection model.

M. Anusha, N. Prashanth, T. Swetha et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.