Skip to content
Conference Open access

Smart Campus Surveillance System: A Multi-Modal AI Approach to Real-Time Threat Detection and Response

2025 · Proceedings of the 1st International Conference on Interdisciplinary Technology & Science Convergence (FusionX Global) · 0 citations · 11 references

TL;DR

The proposed framework unifies the fight and weapon recognition, face identification and contextual interpretation of events into a unified monitoring pipeline, and a novel contribution of this work is the tool calling that allows the VLM to automatically seize important frames and trigger alert protocols in such a way that it will reduce man-in, and response delays.

Abstract

: This study presents the next generation of an intelligent surveillance system for smart campuses based on vision-language models (VLMs) for real-time multimodal detection of threat. The proposed framework unifies the fight and weapon recognition, face identification and contextual interpretation of events into a unified monitoring pipeline. A novel contribution of this work is the tool calling that allows the VLM to automatically seize important frames and trigger alert protocols in such a way that it will reduce man-in, and response delays. The system is deployed on edge devices to balance between computational efficiency and real-time performance and then a centralized surveillance dashboard is used to provide actionable insights by consolidating all the alerts and detections in the surveillance system. Preliminary evaluations show high detection accuracy and low latency, which adds to the prospects of using VLM-driven surveillance in educational environments. Beyond the technical validation, the paper discusses ethical challenges, hardware limitations, and pathways for the easy deployment on a scalable basis ultimately aligning with the SDG 16. This research contributes to the development of proactive and autonomous safety mechanisms by integrating the computer vision, language-based reasoning, and edge AI technologies in an integrated surveillance architecture.

Read PDF

Similar papers

Preprint Aug 2026

City Sentinel: A Unified AI-Based Smart Surveillance Framework for Real-Time Multi-Threat Detection Using Deep Learning

The results demonstrate that a modular, open-source, multi-model architecture can provide broad surveillance coverage, cloud-based auditability, and flexibility for adding new detection capabilities while maintaining practical real-time performance.

Hanan Syed Shabir, Noor Fatima, Safia Baloch et al. · 0 citations
Open access Jul 2026

Integration of CCTV and IoT Sensors for Context-aware Intelligent Surveillance Systems

The growing demand for smarter and more dependable surveillance systems has driven the integration of advanced technologies capable of providing real-time situational awareness and accurate anomaly detection. Conventional Closed-Circuit Television (CCTV) systems are typically constrained by their reliance on manual monitoring and lack of contextual insight, leading to slow response times and high false-alarm rates. This paper presents the design and simulation-based implementation of a context-aware intelligent surveillance system that fuses CCTV video streams with Internet of Things (IoT) sensor data to enhance detection accuracy and system responsiveness, specifically for indoor fire detection. The system employs a Variational Autoencoder (VAE) trained on multimodal data to learn normal behavioural patterns and identify anomalies through reconstruction-error analysis. Training and testing were conducted using simulated sensor data and the MmodalFire dataset, which provides synchronised video and environmental sensing data across diverse indoor environments. The system was implemented within a MATLAB–Python co-simulation framework, enabling effective modelling of both hardware behaviour and software intelligence without physical deployment. Experimental results demonstrate strong performance, with an accuracy of 96.8%, precision of 95.4%, recall of 97.1%, and a low false-alarm rate of 3.2%, representing a significant improvement over a traditional CCTV-only system. The system also exhibits low detection latency, high robustness, and sensitivity under varying environmental conditions, including smoke interference and low visibility. These findings indicate the potential of multimodal data fusion and unsupervised learning to improve surveillance intelligence.

O. M. Nwakeze, Ugoji Frank Godric Chidubem, Nwafor Anthony Chigozie et al. · 0 citations
Open access Aug 2026

REAL-TIME WEAPON AND THREAT DETECTION USING YOLOV12 WITH MULTI-SENSOR FUSION FOR ENHANCED SURVEILLANCE SYSTEMS

This study proposes an Advanced Surveillance Framework that makes use of YOLOv10, a next-generation real-time object detection algorithm that greatly outperforms conventional single-sensor approaches in precision, recall, and real-time responsiveness.

Sadiya Begum, Lubna Nausheen, Ruqiya Fatima · 0 citations
Open access Jul 2026

AI-Powered Real-Time Accident Detection and Emergency Response System with Vehicle Forensic Analysis

Road accidents remain one of the leading causes of injury and death worldwide, and delays in detecting and reporting them significantly increase the risk of severe outcomes. Conventional CCTV-based surveillance depends heavily on human operators, making continuous, error-free monitoring of multiple video feeds impractical. This paper presents an AI-Powered Real-Time Accident Detection and Emergency Response System with Vehicle Forensic Analysis that continuously monitors live CCTV streams using YOLOv11 for vehicle detection and DeepSORT for multi-object tracking to identify collisions automatically. Upon detecting an accident, the system assesses its severity, stores the event in a relational database, and instantly dispatches alerts through SMS, email, voice call, and a monitoring dashboard. A dedicated forensic module then reconstructs the incident by extracting collision frames, recognizing number plates through Optical Character Recognition (OCR), estimating vehicle speed, and retrieving owner records, culminating in an automatically generated digital forensic PDF report for use by police, insurers, and legal authorities. Experimental evaluation across varied traffic and lighting conditions confirms reliable accident detection, fast alert dispatch, and consistent forensic report generation, demonstrating the system's potential to shorten emergency response times and streamline post-accident investigation

Vidya M N, Prajwal Raj V, Dr Manjunath B · 0 citations
Conference Jul 2026

SmartGuard: A Deep Learning and IoT-Enabled Smart Surveillance Framework for Data-Driven Community Safety and Threat Prevention

Conventional closed-circuit television (CCTV) systems record crime rather than prevent it. This paper presents SmartGuard, a smart surveillance framework that combines a fine-tuned YOLOv8-nano deep learning model with an IoT alert pipeline to detect masked or disguised individuals in real time and notify residents before harm occurs. Running on a Raspberry Pi 4 edge device, the system achieves a masked-individual detection mAP50 of 0.995 with sub-1.3 second end-to-end alert latency via Firebase Cloud Messaging. A pilot survey of 32 University of East London participants returned an overall approval mean of 3.88/5. The system avoids facial recognition entirely, operating within UK GDPR constraints. Results demonstrate that effective, low-cost, privacy-conscious residential surveillance is technically feasible on affordable edge hardware.

Md Mahmudur Rahman, Mahmud Yusuf Ahmed, Kazi Tansen et al. · 0 citations
Conference Jul 2026

YOLOv8-Powered Intelligent Surveillance: An Integrated Real-Time Framework for Crowd Management, Crime Prevention, and Workplace Safety Monitoring using AI and ML

The evolving complexity of urban environments and the effectiveness of traditional CCTV surveillance is making it increasingly difficult to ensure public safety, effective crowd management, crime prevention, and workplace security solutions. However, the traditional approach to surveillance is largely manual, leading to late reactions, missed events, and scalability issues. The intelligent surveillance system based on the YOLOv8 object detection algorithm is designed to enhance workplace safety, prevent crimes, manage crowds, and achieve face recognition in real time within a single AI platform. This paper introduces the concept of an intelligent surveillance framework that combines real-time face recognition, workplace safety monitoring, crowd management, and crime prevention through the use of YOLOv8 object detection algorithms within a single AI-driven solution. The framework employs the YOLOv8 algorithm for object detection, identifying people, weapons, suspicious activities, abandoned objects, and workplace safety violations, and issuing automatic alerts to facilitate swift decision-making. The proposed framework provides an integrated platform of multiple surveillance functionalities as opposed to the existing surveillance systems, where each surveillance task is monitored separately, which provides overall situational awareness using the existing CCTV. The model was trained with surveillance images annotated and tested with Precision, Recall, F1-score, Accuracy, and mAP@0.5. An overall detection accuracy of 92.4%, a precision of 92.4%, a recall of 89.7%, an F1-score of 91.0%, and an mAP@0.5 of 93.2% have been achieved during experimental evaluation. Moreover, the framework's average inference latency is 18ms per frame, which guarantees that it can be used in real-time surveillance applications without compromising the accuracy of its detection results when deployed in various surveillance environments. The proposed system is versatile and feasible for implementation in smart city systems, transportation hubs, industrial production sites, and various organizational environments, and can enable smart surveillance by merging multiple security functions into a single framework based on the YOLOv8 object detection model.

M. Anusha, N. Prashanth, T. Swetha et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.