Multi-Level Details Recovery: A Reinforced-Transformer for Person Re-Identification
Abstract
Person re-identification (Re-ID) aims to retrieve the same person across non-overlapping cameras. Despite recent progress, Re-ID remains highly challenging due to high inter-class similarity in large-scale datasets and drastic intra-class variations in cross-platform scenarios (e.g., drones and wearable cameras). While Transformers have been introduced to Re-ID for their superior long-range dependency modeling and robustness against global appearance morphology changes, they often struggle to capture fine-grained discriminative local features due to the inherent properties of the self-attention mechanism. To address this, we propose a Reinforced-Transformer (RT) architecture designed to recover and reinforce these critical local details. Specifically, we develop a Detail Retention Module (DRM) to preserve salient local information while enhancing the interaction between local and global features. Building on this, a Multi-Detail Recovery Module (MDRM) is introduced to progressively restore local features from both coarse-to-fine and fine-to-coarse perspectives. Extensive experiments on four large-scale benchmarks and a mixed aerial-ground dataset demonstrate that our method achieves state-of-the-art (SOTA) or highly competitive performance across different benchmarks.