Skip to content

Author

Kyungyong Chung

1 paper indexed here

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

Open access Sep 2026

Multi-Object Tracking and Spatio-Temporal Graph-Based Relational Pattern Mining

This paper addresses the task of binary violence classification in short surveillance video clips, i.e., deciding whether a given clip contains violent interactions (Fight) or not (Non-Fight). Although this task is often discussed in the broader context of video-based abnormal behavior detection driven by smart cities, Closed-Circuit Television (CCTV) networks, and intelligent surveillance systems, existing methods for violence classification largely rely on appearance features from single frames or global motion cues across whole scenes and therefore fail to adequately capture the interactions among multiple objects and the structural changes in their relationships that arise in real surveillance environments. In particular, violent behavior is rarely defined by a specific pose or a single moment; rather, it emerges as a cumulative process in which relational changes—such as inter-person approach, distance variation, collision, and repeated contact—unfold over time. Detecting such behavior accurately therefore requires an approach that can analyze inter-object relationships in a spatiotemporal manner. To this end, this paper proposes a violence detection method that combines multi-object tracking with spatiotemporal graph-based relational pattern mining. The proposed method first detects and tracks person objects using YOLO and DeepSORT, and extracts time-series features—including position, velocity, pose, and inter-object distance variation—to construct a spatiotemporal graph. Relational event sequences are then generated from the edge features of the graph, and class-representative relational patterns are automatically extracted based on discriminative power through PrefixSpan-based frequent sequential pattern mining. In parallel, the spatiotemporal graph is fed into a Spatial Temporal Graph Convolutional Network (ST-GCN) to learn the structural relationships among objects and their temporal evolution. Finally, the pattern-matching score and the ST-GCN classification score are combined to classify each input video as either violent or non-violent. By jointly exploiting interpretable relational pattern information and graph-based structural learning, the proposed approach compensates for the limitations of appearance-centric anomaly detection and demonstrates its applicability to complex real-world surveillance environments. Performance is evaluated in terms of Accuracy, Precision, Recall, and F1-score, with Recall considered a primary metric to reflect the importance of not missing violent events.

Tae-Yaung Seo, Kyungyong Chung · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.