Skip to content
Open access

Object-Centric Industrial Video Anomaly Detection with Local-Global Representation Learning

Aug 2026 · International journal of information and communication technology trends · 0 citations

TL;DR

A novel framework centered on object-centric video anomaly detection, heavily augmented by a local-global representation learning mechanism, suggesting that integrating structured object interactions into representation learning provides a highly scalable and robust solution for real-world industrial monitoring.

Abstract

Industrial video anomaly detection is a critical component of modern smart manufacturing and industrial surveillance, aiming to automatically identify deviations from normal operational patterns. Traditional methods relying on frame-level or pixel-level feature extraction often struggle with the complex, dynamic, and heavily occluded environments typical of industrial settings. This paper proposes a novel framework centered on object-centric video anomaly detection, heavily augmented by a local-global representation learning mechanism. By shifting the analytical focus from the entire image frame to specific objects of interest, such as machinery components, manufactured goods, and human operators, the proposed method isolates highly relevant features while mitigating the impact of background noise and illumination variations. The framework utilizes a robust tracking-by-detection paradigm to construct spatio-temporal object tubes, which are subsequently processed to extract localized representations capturing appearance and motion dynamics. Concurrently, a global representation module employs attention mechanisms to model the complex interactions between multiple objects and their contextual environment. The integration of these local and global streams ensures a comprehensive understanding of the industrial scene, allowing for the precise localization and classification of anomalous events. Extensive evaluations on multiple large-scale industrial datasets demonstrate that the proposed object-centric framework significantly outperforms existing state-of-the-art approaches in both anomaly detection accuracy and computational efficiency. The findings suggest that integrating structured object interactions into representation learning provides a highly scalable and robust solution for real-world industrial monitoring. 

Read PDF

Similar papers

Jul 2026

O-VAD: Industrial Video Anomaly Detection through Object-Centric Tracking and Reasoning

Industrial Video Anomaly Detection (IVAD) aims to identify anomalous objects and events in an industrial process, which is crucial for modern manufacturing and quality control systems. Existing VLM-based anomaly reasoning methods are capable of detecting open-ended anomalies in general domains. However, their performance declines in industrial settings characterized by intricate object transformations, strict physics, and procedural constraints. To tackle the complexity of such interaction-intensive detection, we introduce a training-free agentic framework for anomaly detection free of domain-specific knowledge, emphasizing object state evolution like humans inspectors. It is designed to track spatial-temporal dynamics and underlying transformations of detected objects over time, and then reason over the object-wise temporal state trajectories to identify abnormal objects in grounded frames. Our method overcomes limitations of prior approaches that rely on retraining on normal clips or injecting domain knowledge as context for test-time inference. Extensive experiments on three IVAD datasets demonstrate that our method outperforms frontier VLMs, agentic frameworks, and traditional VAD methods fine-tuned on the respective datasets, while providing interpretable reports over anomaly processes and types.

Mei Yuan, Qi Long, Qifeng Wu et al. · 0 citations
Open access Jul 2026

Enhanced deep learning model for anomaly object detection and tracking from surveillance videos.

An enhanced wolf Crocuta optimization-based deep Bidirectional Long Short-Term Memory (EnWC-DBiLSTM) classifier is proposed using an enhanced wolf Crocuta optimization-based deep Bidirectional Long Short-Term Memory (EnWC-DBiLSTM) classifier for anomaly object detection and tracking.

B. Gayal, S. Patil, D. Meshram et al. · 0 citations
Open access Aug 2026

HIGH-PRECISION AERIAL OBJECT DETECTION MODEL UTILIZING YOLO V10 DEEP NEURAL NETWORK

The installation of a real-time visual tracking system with an active pan-tilt camera for indoor human motion detection is presented, which shows that the inclusion of YOLOv10 significantly improves detection precision and temporal consistency.

Ayman Javid Hussain, Lalitha Saroja Ch, Ruqiya Fatima · 0 citations
Preprint Aug 2026

STEP: Score-Based Temporal Energy for Human Pose Video Anomaly Detection

Skeleton-based Video Anomaly Detection (VAD) offers a robust, privacy-preserving solution for identifying abnormal behaviors. To model the distribution of normal static and moving poses, recent methods train Energy-Based Models (EBMs) via Denoising Score Matching (DSM). However, directly injecting noise, required for training, into raw joint coordinates creates physically impossible poses, and this structural collapse severely worsens as the temporal window expands. To address this, we introduce STEP, a simple framework that utilizes Principal Component Analysis (PCA) to project pose sequences into a compact, whitened PC-space. Learning the data density within this well-behaved PC-space ensures that the injected noise translates into physically plausible variations, which allows the model to process longer video sequences without the performance collapse of raw coordinate baselines. Additionally, to mitigate inherent pose estimation inaccuracies arising from occlusions or motion blur, we integrate a sequence-level weighting mechanism based on the estimator's confidence scores. Operating at real-time computational efficiency, our simple and lightweight framework outperforms the previous skeleton-based state-of-the-art by 12.2% (90.1% AUROC) on the challenging UBnormal dataset and achieves highly competitive results by improving on the ShanghaiTech benchmark.

Jakub Micorek, M. Koziński, Horst Possegger · 0 citations
Conference Aug 2026

Dynamic blur association rule extraction method in real-time visual tracking

A dynamic blur-aware association rule extraction framework is presented for real-time visual tracking in environments characterized by severe and varying motion blur, frequent occlusions, and rapid background changes. The approach innovatively models blur as a spatially and temporally variant process, embedding local blur statistics into compact feature descriptors that are central to the formation of robust association rules. This technique enables the tracking system to selectively suppress unreliable information caused by transient blur. This allows them to distinguish between target dynamics and visual artifacts. The blur-adaptive association module is used to optimize computational efficiency and spatiotemporal accuracy. The module mainly focuses on flow-guided marginal loss and sparse optical flow strategy. The architecture adopts a hybrid CPU/FPGA acceleration scheme to meet the demanding real-time requirements. In the experimental study of UAV urban navigation and robot operation platform, the framework shows high accuracy, strong anti-drift and stability in complex high-speed environment. Comparative analysis and component ablation experiments show that motion-guided optimization, adaptive association and blur modeling have a significant impact on system robustness. These results establish the framework as a significant step forward for reliable, real-time visual tracking in challenging, dynamic settings.

Wen Zhang, Chao Shi · 0 citations
Jul 2026

Integrated Spatial–Temporal Framework for Video Anomaly Detection in Surveillance Systems

Video-based anomaly detection seeks to discover anomalous events, such as crimes, fires, or medical emergencies, by utilizing both spatial and temporal features of video data. Traditional surveillance systems are frequently limited to minimal recording, requiring human analysts for post-event assessment, resulting in delayed responses during crucial occurrences. To address these issues, we present a multi-layered approach to detecting video anomalies that can deal with both temporal and spatial components of video data. The input video is initially obtained from the dataset and undergoes frame conversion. The extracted key frames are then preprocessed for further analysis. To obtain multi-scale spatial characteristics from each frame, the first layer uses a spatial Pyramid pooling network (SPP-Net) along with a convolutional neural network (CNN). These spatial features are then passed to an optimized bi-directional gated recurrent unit (Opt-Bi-GRU) enhanced with Multi-Head Self-Attention (MHSA), which analyzes the temporal dynamics and captures both forward and backward dependencies across frames. Finally, a capsule network (CapsNet) processes the output of the Bi-GRU, identifying complex patterns that may indicate abnormalities over time. The proposed method is implemented using Python. The proposed model performs better than existing methods in terms of F1-score, specificity, sensitivity, accuracy, recall, precision, FPR, and FNR. The proposed model achieves the highest accuracy of 98.2%, 98.87%, and 98.52%, respectively, utilizing the UBI-fights, UCF-crime, and UCSD pedestrian datasets. These results demonstrate that the proposed framework provides an automated, reliable, and effective solution for real-time anomaly detection in surveillance systems.

M. Rao, Priyesh Kumar · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.