Skip to content

CLIP-SGI: A Semantic-Guided and Instance-Consistent Framework for Generalizable Person Re-Identification.

Aug 2026 · IEEE Transactions on Image Processing · Vol PP · 0 citations
Medicine

TL;DR

This work proposes CLIP-SGI, a semantic-guided and instance-consistent framework for generalizable person ReID that combines semantic guidance, domain-aware representation learning, and instance consistency to improve robustness under domain shifts.

Abstract

Generalizable person re-identification (ReID) requires a model trained on labeled source domains to remain discriminative in unseen environments. Although vision-language models provide rich cross-modal priors, the appearance semantics used by CLIP-based ReID are often encoded implicitly in learned prompts and are not explicitly organized into reusable part-level cues. Moreover, conventional identity supervision mainly emphasizes the separation of source identities and makes limited use of the local relationships among visually similar instances. To address these issues, we propose CLIP-SGI, a semantic-guided and instance-consistent framework for generalizable person ReID. First, multiple off-the-shelf vision-language models generate pedestrian descriptions. For each VLM, upper- and lower-body attributes are first voted across images of the same identity and then across the VLM-specific identity labels. Second, we construct an Attribute Prototype Bank (APB) that uses the resulting attributes as region-aware semantic anchors to guide appearance-sensitive feature learning. Third, we introduce a similarity-aware and frequency-normalized soft-label constraint that preserves ground-truth identity supervision while exploiting reliable neighborhood relationships as auxiliary signals. The three-stage training scheme combines semantic guidance, domain-aware representation learning, and instance consistency to improve robustness under domain shifts. Extensive experiments on multiple benchmark datasets demonstrate the effectiveness of the proposed method and its consistent improvements in mean average precision (mAP) and Rank-1 (R1) accuracy.

View source

Similar papers

Preprint Aug 2026

FaLCon: Facet-Anchored Retrieval with Late Consensus for Sim2Real Text-Based Person Anomaly Search

Text-based person anomaly search requires retrieving real-world pedestrian images from detailed natural-language descriptions using models trained primarily on synthetic data. This Sim2Real setting is particularly challenging because visually similar candidates may differ only in subtle actions, object interactions, or appearance attributes, while applying multimodal large language models to the entire gallery is computationally expensive. We propose an anchor-constrained coarse-to-fine retrieval framework that combines global semantic matching with fine-grained verification. First, each query is represented by its original caption, a structured concatenation, and several semantic facets. Heterogeneous vision-language retrievers are then integrated through robust per-query score calibration and soft claim-aware fusion. Full and concatenated captions serve as anchors to preserve candidate recall, whereas appearance, action, and object facets provide bounded corrective evidence. The resulting candidate pool is further refined by a discriminative Qwen3 reranker and two complementary semantic verification modules based on anomaly-aware cloze completion and multi-agent evidence reasoning. Finally, an uncertainty-gated consensus module adaptively reweights the three experts on ambiguous queries. Experiments on the PAB benchmark show that the proposed soft claim-aware retrieval achieves 86.44% mAP@10, substantially outperforming individual retrieval backbones. The complete framework further improves performance to 95.41% mAP@10, 94.44% R@1, and 99.09% R@5. These results demonstrate that preserving strong global retrieval while restricting expensive semantic reasoning to a small candidate pool is effective for fine-grained Sim2Real person anomaly search. Our code will be available on Github.

H. Pham, P. Tran, Thuan Duc Mai et al. · 1 citation
Aug 2026

IRPP: Invariant Representation Learning With Progressive Prototype Refinement for Unsupervised Person Re-Identification

Unsupervised person re-identification (USL-ReID) typically relies on clustering to generate pseudo-labels, but significant cross-view appearance variations often cause images of the same identity to be split into different clusters. Training on such noisy pseudo-labels severely degrades the learned representations. Therefore, learning robust view-invariant features is paramount. Data augmentation provides a direct way to enhance invariance, yet its trade-offs in USL-ReID remain under-explored: weak augmentations usually preserve identity semantics but lack diversity, whereas strong augmentations provide richer appearance diversity at the cost of partially corrupting identity-consistent semantic cues. To address this challenge, we propose Invariant Representation learning with Progressive Prototype Refinement (IRPP), a unified framework that learns invariant and discriminative features from noisy pseudo-labels. IRPP consists of three synergistic components. First, an Augmented Dual-Contrastive Learning (ADCL) module performs dataset-level prototype-guided invariant learning by contrasting weakly and strongly augmented views against cluster-derived prototypes. Second, an Alignment and Uniformity Learning (AUL) module regularizes the mini-batch-level weak–strong feature geometry, leading to more stable feature distributions under data augmentation. Third, a Progressive Prototype Refinement (PPR) mechanism progressively optimizes cluster centroids into cleaner prototypes, thereby mitigating the influence of noisy pseudo-labels and further strengthening invariant representation learning. This closed-loop design enables prototype-guided contrastive learning, weak–strong regularization, and prototype refinement to mutually reinforce each other. Extensive experiments on standard USL-ReID benchmarks demonstrate that IRPP achieves state-of-the-art performance with a simple and efficient training pipeline. Code is available at https://github.com/Trangle12/IRPP

Xuan Tan, Qixian Zhang, Ding Qi et al. · 0 citations
Jul 2026

Multi-Level Semantic-Guided Framework for Cloth-Changing Person Re-Identification

A Multi-level Semantic-Guided (MSG) framework that integrates contextual and fine-grained visual information to eliminate clothing variance across both conceptual and pixel dimensions is proposed.

Shijuan Huang, Hefei Ling, Zongyi Li et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.