Skip to content

VLA-ReID: Video-Level Association for Re-Identification in Multi-Object Tracking with Highly Similar Objects

Jul 2026 · arXiv.org · Vol abs/2607.17157 · 0 citations · 48 references
Computer Science

TL;DR

Video-Level Association re-ID (VLA-ReID), which reformulates re-ID as video-level association modeling, and uses aggregated historical trajectory features as queries and all current-frame detections as candidates as candidates, enabling direct optimization of their global association at each frame.

Abstract

Multi-object tracking (MOT) aims to localize multiple objects in videos while preserving their identities over time. Long-term identity preservation remains difficult when objects are small, densely distributed, and highly similar in appearance, as in bee swarm scenes. Existing trackers rely on re-identification (re-ID) models trained through single-instance assignment (instance-level querying). At inference, however, MOT requires global assignment between multiple trajectories and detections, corresponding to video-level querying. This training-inference mismatch can cause identity switches among visually similar objects. Existing approaches also often require substantial additional annotations to enhance appearance discrimination. We propose Video-Level Association re-ID (VLA-ReID), which reformulates re-ID as video-level association modeling. It uses aggregated historical trajectory features as queries and all current-frame detections as candidates, enabling direct optimization of their global association at each frame. In addition, Frame-Common Appearance Estimation (FCAE) estimates a common appearance direction from current-frame detections, while Common-Appearance Suppression (CAS) removes the corresponding component along this direction from trajectory and detection features. This amplifies discriminative differences among highly similar objects without additional annotations. Experiments on BEE24 show that VLA-ReID improves HOTA by 1.1, MOTA by 0.3, AssR by 2.6, AssA by 0.7, and IDF1 by 0.8 over state-of-the-art trackers, while reducing identity switches by 28%. These results demonstrate the effectiveness of video-level re-ID modeling for appearance-based association in MOT.

View source

Similar papers

Conference Aug 2026

PrismTrack: Perspective-Aware Multi-Cue Association for Robust Multi-Object Tracking

Multi-Object Tracking (MOT) remains challenging due to object occlusion, complex motions, and detection unreliability in crowded scenarios. We propose an enhanced MOT framework integrating and optimizing state-of-the-art components, specifically Improved Detection Confidence Boost (IDCBoost) and Track-Perspective-Based...

Trung Nghia Huynh, Chi Nhan Huynh, Jia-Ching Wang et al. · 0 citations
Review Sep 2026

Tracking-by-detection in Multi-object Tracking: Survey and Experiments

Multi-object tracking (MOT) is an essential computer vision task that simultaneously tracks multiple objects in video sequences, with various applications in surveillance, autonomous navigation, and human-computer interaction. The tracking-by-detection (TBD) paradigm, which combines object detection with temporal assoc...

Yu-Jin Yang, Kyujin Shim, Kangwook Ko et al. · 0 citations
Preprint Sep 2026

Online Multi-Camera 3D Tracking via ID Prediction over Recurrent Sparse Queries

Online multi camera 3D tracking must maintain scene global identities across synchronized views, yet query-based trackers carry these identities only implicitly in the instance bank, where they fragment upon query interruption. We present an online architecture that recovers association accuracy by predicting IDs expli...

Pragyan Shrestha, H. Nakayama, Atom Scott · 1 citation
#small language model Open access Sep 2026

Reliability-Aware Semantic Gating with Retrieval-Augmented Association for Online Multi-Object Tracking Under Industrial Low-Altitude Proxy Conditions

Industrial low-altitude multi-object tracking is challenged by overhead viewpoints, dense small targets, mutual occlusion, homogeneous appearances, and domain-shifted backgrounds. Vision-language pre-trained models provide useful semantic priors, but their category-level representations can mislead instance-level assoc...

Rong-Zuo Guo, Yu-Hang He · 0 citations
#artificial intelligence Preprint Sep 2026

Tracking the Unseen: An Occlusion-Robust Framework for Target Tracking Under Full and Long-Term Occlusion

Real-time multi-object tracking systems remain highly vulnerable to full and long-term occlusion, where targets temporarily or completely disappear from the camera's field of view. Conventional trackers may terminate trajectories prematurely, resulting in identity loss and reduced situational awareness in applications...

Mais.M Mohammed, Sharifa Mohammed, Hanan Awadh et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.