Jul 2026· ACM Transactions on Multimedia Computing, Communications, and Applications (TOMCCAP)· 0 citations· 48 references
TL;DR
A Multi-level Semantic-Guided (MSG) framework that integrates contextual and fine-grained visual information to eliminate clothing variance across both conceptual and pixel dimensions is proposed.
Abstract
Cloth-changing person re-identification (CC-ReID) aims to match persons who change clothing across multiple surveillance cameras. Recent approaches strive to extract clothing-agnostic features by utilizing biological information, including skeleton, texture, gait, and 3D data. However, most methods rely on additional features at a single aspect, leading to a lack of comprehensive understanding of concepts and semantics. This limitation introduces biases and restricts both accuracy and functionality, thereby diminishing their effectiveness in handling variations in clothing. To alleviate this problem, we propose a Multi-level Semantic-Guided (MSG) framework that integrates contextual and fine-grained visual information to eliminate clothing variance across both conceptual and pixel dimensions. This innovative solution consists of two key components: the Contextual Semantic Guidance (CSG) module adeptly utilizes textual features from clothing descriptions to decouple clothing concepts at a higher semantic level. In contrast, the Low-Level Robust Feature Disentanglement (LRD) module meticulously analyzes images featuring clothing to disentangle texture information at the granular pixel level, and the integration of a momentum update mechanism significantly bolsters the model’s robustness. The two approaches collaboratively eliminate clothing information by interacting across different semantic levels. Extensive experiments demonstrate the effectiveness of our method, achieving new state-of-the-art performance on several popular CC-ReID benchmarks. Our code will be available on GitHub at https://github.com/ShijuanHuang/MSG.
Clothes-changing person re-identification (CC-ReID) is a challenging task with great practical value. Existing methods attempt to decouple clothes-independent features from the RGB modality alone and suffer from spatial redundancy. In this paper, we propose a novel Attribute-aware Prompt Learning (APL) framework for CC-ReID that comprises a frequency-aware visual prompt generator (FVPG) and a set of attribute-aware text prompts (ATP). Specifically, FVPG computes high-frequency components of images to provide additional body-shape information at low computational cost. ATP employs a CLIP model to learn clothes-independent features with the designed attribute-aware text prompts via vision-language contrastive learning. With a proposed two-stage training strategy, the APL framework can capture clothes-independent features, effectively improving the performance of CC-ReID. Extensive experiments demonstrate that the proposed APL can achieve state-of-the-art performance on three widely used clothes-changing person ReID benchmarks.
Person search in real-world surveillance and robotic perception often fails when facial cues are unreliable due to low resolution, motion blur, or occlusion. We propose a clothing-centric person search framework that represents people using appearance attributes (e.g., garment type, dominant colors, accessories) instead of biometrics. Given person tracklets, the system samples frames, extracts clothing cues, and generates schema-constrained structured descriptions using a large language model, enabling consistent semantic indexing. Retrieval is performed with attribute-aware semantic search over these descriptions and compared against appearance-based baselines in controlled multi-camera experiments. A full-dataset analysis of the generated descriptions highlights common failure modes under real-world conditions, including color ambiguity under lighting changes (e.g., black vs dark), posture hallucinations under occlusion (e.g., standing behind a table described as sitting), and occasional prompt deviations that introduce spurious attributes. Runtime results show that detection is lightweight (about 30 ms per frame), while semantic extraction dominates; with parallel workers, the full pipeline processes roughly 0.25-0.34 s of computation per second of video, supporting practical deployment.
Diana Souza, Khansa Rekik, Rainer Müller et al.· International Conference on...· 0 citations
This work proposes CLIP-SGI, a semantic-guided and instance-consistent framework for generalizable person ReID that combines semantic guidance, domain-aware representation learning, and instance consistency to improve robustness under domain shifts.
Dai-Xin Liu, Yu Yang, Linlin Tang et al.· IEEE Transactions on Image P...· 0 citations
Universal person re-identification (ReID) aims to retrieve pedestrian identities across diverse real-world scenarios, including severe occlusions, clothing changes, and cross-modality shifts, within a unified model. However, existing 2D representations fundamentally struggle with spatial ambiguities due to a lack of depth and topological awareness, while naively introducing monocular 3D priors often causes severe negative transfer due to geometric estimation noise under extreme visual degradation. To safely harness the clothing-invariant and canonical structural properties of 3D geometry, we propose UniGeo, a Universal Monocular 3D-Enhanced ReID framework driven by a Consistency-Aware Reliability Gate and Dual-Stream Residual Fusion. Specifically, the processing of 3D information is strategically decoupled into geometric extraction and dynamic utilization. To provide pure structural compensation, we project monocular 3D parameters into kinematic joint representations, explicitly capturing instance-level geometric topology to resolve appearance-based ambiguities. To robustly incorporate these cues without perturbing the reliable 2D feature space, we isolate the 3D prior as a late-stage structural residual; modulated by the consistency-aware gate, this mechanism adaptively filters geometric noise and enables controlled fallback to the pure 2D baseline. Extensive experiments show that our method improves challenging, structure-sensitive scenarios while preserving competitive performance on clean domains. Code is available at https://github.com/BohanSu/UniGeo.
Bohan Su, Jiashuo Wang, Fangyi Liu et al.· 0 citations
A disentanglement method based on Mutual Information Minimization is introduced to minimize statistical dependence between modality-shared and modality-specific features from a probability distribution perspective and a Local Discriminative Attention Module is designed to adaptively focus on highly informative body parts such as head-shoulder ratio and torso patterns.
Xiaokai Liu, Fangqing Zhou, Qian Song et al.· Advances in Engineering Tech...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.