Stratified Consistency Learning for Text-to-Image Person Re-identification
Abstract
This work addresses critical challenges in text-to-image person re-identification (TIReID), including inconsistent correspondences across samples in human-annotated datasets and insufficient exploration of fine-grained consistency information between sample pairs. To mitigate the adverse effects of inter-modal inconsistencies at the global level, a Stratified Consistency-Perceptive Unit (SCU) is designed. It constructs stratified-granularity features from local feature sequences and corresponding cross-modal correlation weights derived from CLIP. Cross-modal losses among same-level granularities are modeled as consistency signals, which are then aggregated across semantic levels to accurately assess sample-level consistency. To address fine-grained semantic inconsistencies, an Implicit Consistency Learning Unit (ICU) is designed. It applies masked modeling to both low-level (word-level) and high-level (phrase-level) representations, integrating contextual semantics from diverse structures to capture fine-grained consistency across samples. Extensive experiments conducted on the CUHK-PEDES, ICFG-PEDES, and RSTPReid datasets, comprising 40,206, 54,522, and 20,505 images respectively, comprehensively validate the effectiveness and robustness of the proposed framework, achieving a 1.4% improvement in mean Average Precision over existing state-of-the-art methods.