Skip to content
Open access

An end-to-end multi-modal pipeline for person search in unconstrained CCTV environments

Sep 2026 · Scientific Reports · 0 citations

Abstract

In real-world surveillance, situations arise in which two people wear nearly identical clothing or in which identifying features are obscured by heavy occlusions and shifting poses. In these environments, traditional uni-modal systems that rely on static appearance do not perform well and often produce false matches. This paper presents a multi-modal Re-ID framework that combines appearance, facial features, and gait to accurately perform this task. In the first stage, robust person tracklets are extracted from video using a YOLOv8 detector and a DeepSORT tracker. A multi-modal feature extractor is then applied to each person crop through three parallel streams. A custom GaitSet-inspired network produces a 256-dimensional gait embedding from RGB tracklet frames. An OSNet backbone is used to get a 1024-dimensional omni-scale appearance embedding. For facial biometrics, a Multi-Task Cascaded Convolutional Network (MTCNN) is used to explicitly detect the face and five-point landmarks followed by a facial landmark regression model. Each crop is aligned to a landmark and the aligned face is sent to FaceNet for computing the 128 dimensional facial embedding. In the last step, an Ensemble Concatenation module merges these three normalised embeddings into a single $$\mathbb {R}^{1408}$$ identity descriptor ( $$1024 + 256 + 128$$ ). The discriminability is significantly improved through the combination of facial features and temporal motion cues. Furthermore, the fusion strategy prevents any single modality from causing identity leakage, thereby providing a strong defense against appearance-based identity errors.

Read PDF

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.