Visible-infrared person re-identification (VI-ReID) is an important technique for around-the-clock person matching, and its primary challenge arises from substantial cross-modal discrepancies. To address this challenge, we propose a Decoupled Information-Guided Cross-Modal Alignment (DIGCA) framework organized into three stages: multi-scale feature modeling, partial functional decoupling, and guided cross-modal alignment. First, a Hierarchical Context Extractor (HCE) aggregates multi-scale contextual information through dilated convolutions and residual connections to enrich the initial identity representation. Second, a Multi-Branch Unified Encoder (MBUE) organizes complementary feature streams through parallel global-relation and local-spatial modeling. The resulting representations exhibit different information tendencies and promote partial functional decoupling between identity-related and modality-related information. Finally, the Decoupled Information-Guided Cross-Modal Alignment (DIG-CMA) module refines the two streams with channel and spatial attention and integrates them through cross-attention. Under the joint optimization objective, the resulting representations support cross-modal alignment. Experiments on SYSU-MM01, RegDB, and LLCM show that DIGCA achieves competitive overall performance compared with recent VI-ReID methods, providing empirical support for the effectiveness of the proposed decoupling-guided alignment strategy in cross-modal identity matching.
Abstract. Existing single-object trackers often struggle to fully exploit temporal information and distinct feature representations, limiting their robustness in complex scenarios. Specifically, current approaches face challenges in (1) capturing channel-wise discriminative features during target deformation, (2) adapting to sudden appearance changes due to reliance on static memory banks, and (3) mitigating background interference within transformer attention mechanisms. To address these issues, we propose STFTrack, a proposed tracking framework integrating enhanced spatiotemporal features. First, we construct a semantic–spatial–instance attention module, which refines target representation via cascaded channel calibration and instance uncertainty modeling. Second, a gated attention fusion module is introduced to adaptively aggregate temporal and spatial cues. Third, we design an improved spatiotemporal decoder equipped with learnable autoregressive queries to establish a robust cross-frame state propagation mechanism. Extensive experiments on five benchmarks, including LaSOT, GOT-10k, and TrackingNet, demonstrate the superiority of STFTrack. For instance, it achieves an area under the curve score of 72.8% on the LaSOT benchmark. Qualitative results further confirm that STFTrack effectively suppresses background clutter and maintains accurate tracking under challenging conditions including deformation, fast motion, and occlusion.
Changhao Zhou, Xun Duan, Guangqian Kong et al.· Journal of Electronic Imagin...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.