Segmentation-Guided Scene Perturbation for Passage-Based Top-View Person Re-Identification
Abstract
Passage-based top-view person re-identification aims to match individuals across short overhead walking videos. This privacy-preserving setting is challenging because overhead cameras suppress facial and body-part cues, compress pedestrian appearance, and make models vulnerable to scene shortcuts from floors, ramps, and doorways. Existing ReID pipelines are mainly designed for standard pedestrian views, where clothing details and part-level structures are more reliable. We address this gap with a compact scene-agnostic RGB framework that targets shortcut suppression rather than architecture-heavy part modeling. The core idea is segmentation-guided human-aware augmentation, which uses pedestrian masks during training to weaken background reliance and encourage person-centered descriptor learning. A single-branch Vision Transformer is then trained with Circle-based discriminative supervision, and descriptor-consistency frame selection is applied before passage-level aggregation at inference. Under the official ICPR 2026 TVRID RGB-track protocol, the proposed method ranks first across the public and private evaluation scenarios and substantially outperforms the organizer-provided RGB baseline. Additional evaluation on Market-1501 achieves 94.7% Rank-1 and 87.2% mAP, showing that the compact training recipe remains competitive beyond the target top-view setting.