Skip to content
Conference

Edge-AI Enabled Real-Time Multi-Camera Casualty Detection and Tracking for Military IoT

Jul 2026 · International Conference on Ubiquitous and Future Networks · pp. 1146-1151 · 0 citations · 24 references

Abstract

The advancement of military internet of things (IoT) surveillance demands real-time casualty detection across distributed camera networks under diverse environmental conditions. Traditional surveillance systems suffer from high latency, bandwidth inefficiency and unreliable cross-camera identity tracking, indicating the need for advanced detection and tracking systems. This study presents an edge-based multicamera casualty detection and tracking system for military IoT networks. The proposed framework, called CCTV-TrackNet, integrates lightweight AI on distributed camera nodes built on Raspberry Pi Zero2W hardware with high-resolution cameras and sensors such as GPS and IMU. At each node, YOLOv12n performs real-time person and casualty detection, DeepSORT maintains short-term continuity, and OSNet-based Re-ID extracts appearance embeddings for cross-camera association. A central server aggregates metadata for global identification, visualizes trajectories, and generates real-time alerts. Experimental results show 90.77% detection accuracy, 37% and 36% reduction in false cases, 30.3 FPS edge performance, and 84.20% cross-camera ID consistency.

View source

Similar papers

Preprint Aug 2026

City Sentinel: A Unified AI-Based Smart Surveillance Framework for Real-Time Multi-Threat Detection Using Deep Learning

The results demonstrate that a modular, open-source, multi-model architecture can provide broad surveillance coverage, cloud-based auditability, and flexibility for adding new detection capabilities while maintaining practical real-time performance.

Hanan Syed Shabir, Noor Fatima, Safia Baloch et al. · 0 citations
Conference Jul 2026

Skeleton-based Anomaly Detection for CCTV Surveillance: A Real-Time Framework with Emergency Escalation

Conventional video surveillance based on pixel-level deep-learning models is resource hungry, processes gigabytes of video material, and retains biometric identifying data. This paper describes a lightweight, privacy-sensitive alternative that uses skeletal pose estimation to replace pixel-based processing. We only process 33 coordinates for body joints using MediaPipe BlazePose, compressing the data by 2000× and discarding all visual identity data at the outset. The innovation is a hybrid dual- layer classifier that combines a Random Forest classifier trained on spatiotemporal features and deterministic rules from physics (velocity thresholding and kinematic plausibility), these rules help handle ambiguous cases when the Random Forest classifier is uncertain. Testing on 1145 frames of normal, suspicious, and dangerous activities achieves an overall accuracy of 87.42% with 0.52 precision and 0.89 recall for dangerous activities. The system has under 15 ms inference time on standard laptop CPUs and under 35 ms on Raspberry Pi 4, supporting real-time edge computing without GPUs. The framework reduces computation compared to CNN-based approaches and can run on $50 edge devices. Experiments show that skeletal geometric properties are sufficient for behavior classification without relying on appearance-based biometric features.

S. L. Jany Shabu, P. Asha, P.Asmitha Priyaa et al. · 0 citations
Conference Jul 2026

Data Fusion of Cameras attached on Dual UAVs for Real-Time Human Detection and Localization

Unmanned aerial vehicles (UAVs) are increasingly used in search and rescue operations due to their rapid deployment and wide-area coverage. However, many existing UAV-based systems focus mainly on victim detection and do not provide accurate geographic coordinates, which limits their usefulness in real rescue missions. In addition, detecting victims from aerial imagery remains challenging when targets appear as very small objects. This paper proposes a real-time dual UAV system for human detection and localization that integrates deep learning and geometric triangulation. Two UAVs equipped with cameras capture synchronized visual and telemetry data, which are transmitted to a ground processing server. A deep learning model based on YOLO is used to detect victims and guide semi-automatic camera alignment, while geometric triangulation and telemetry data from two independent UAVs are combined to estimate the victim’s GPS coordinates. The proposed system was evaluated through real-world field experiments. Experimental results show that the system achieves an average localization error of approximately 3.6 meters at an observation distance of about 150 meters, while maintaining real-time processing performance. These results demonstrate the feasibility of combining multi-UAV vision and geometric modeling to improve the effectiveness of UAV-assisted search and rescue operations.

T. Do, Tat-Dat Nguyen, B. Nguyễn et al. · 0 citations
Open access Jul 2026

Real-Time Spatio-Temporal Deepfake Detection for Live Biometric Authentication via EfficientNet-GRU

Deepfake technology poses a critical threat to live video conferencing and biometric authentication. Existing detection models are either purely spatial—rendering them vulnerable to video compression—or rely on computationally heavy 3D-CNNs incompatible with strict real-time CPU latency constraints. We propose a highly optimized Two-Stream Spatio-Temporal architecture specifically designed for zero-latency live video evaluation. The framework extracts fine-grained spatial artifacts using a lightweight EfficientNet-B0 backbone, while a unidirectional Gated Recurrent Unit (GRU) models frame-to-frame physiological inconsistencies. Evaluated on a diverse 3,000-video corpus from FaceForensics++, Celeb-DF, and DFDC, the model achieved 98.21 cross-dataset accuracy and a 0.9978 AUC. Crucially, CPU inference requires only 125.63ms per 16-frame sequence—well below the 533ms threshold of a standard 30 fps camera—guaranteeing seamless, real-time overlay detection. Finally, the network's decision boundaries are mathematically validated using Explainable AI (XAI) activation maps and t- SNE clustering.

Saurabh Jha, Akash Sanghi, Pragati Upadhyay et al. · 0 citations
Open access Aug 2026

An Edge-Deployable Lightweight UAV Detection and Net-Capture System Based on NCDet-YOLO

The growing frequency of unauthorized UAV activities has increased the demand for real-time perception and rapid response on resource-constrained edge devices. This study proposes an edge-deployable UAV detection and net-capture system based on Net-Capture Detection YOLO (NCDet-YOLO). Developed from YOLOv8n, NCDet-YOLO incorporates C2f_Faster, SPD_Conv, EMA, and a lightweight three-scale detection head, with CrossKD used to compensate for accuracy loss caused by structural compression. The dataset contains 6615 images and was divided into 5292 training and 1323 validation images. The self-collected data include DJI Phantom 4 and DJI Inspire 2 UAVs observed at approximately 4–30 m under different daytime backgrounds. NCDet-YOLO achieves an mAP50–95 of 0.6504 with 1.55 M parameters and 4.1 GFLOPs. On a Jetson Orin NX Super under the 15 W power mode, it achieves 31.53 FPS, representing a 31.67% increase over YOLOv8n. The detector is further integrated with target alignment, distance determination, trigger control, and net-capture execution. In 10 real-platform trials, 8 captures were successful, corresponding to an 80.0% success rate, with one false-trigger event and an end-to-end latency from target detection to net-capture firing of approximately 400 ms.

Jinting Ye, Jie Lang, Kefei Liao et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.