Jul 2026· IEEE Transactions on Image Processing· Vol 35, pp. 8206-8218· 0 citations· 73 references
MedicineComputer Science
Abstract
Referring Video Object Segmentation (RVOS) aims to segment the target objects specified in human instructions. Previous approaches typically rely on explicit human instructions that contain target categories or salient appearance descriptions. These approaches tend to fail when the instructions require temporal video understanding and complex relational reasoning. In this work, we present RViSeg, a reasoning-centric video object segmentation model that leverages the reasoning capability of Multi-modal Large Language Models (MLLM) to handle complex queries. The primary challenge lies in enabling MLLM to perform efficient pixel-level video perception. To tackle this challenge, we introduce a novel Spatial Token Merge (STM) module that consolidates lengthy video tokens into compact region-level clusters, while preserving essential spatial details. This structured representation enables MLLM to infer user intention by interleaving spatial and temporal visual information. Furthermore, we propose a Query-based Target Retrieval (QTR) module that utilizes learnable tokens as the target identity for mask prediction. By propagating these instance-specific tokens both intra-clip and inter-clip, our RViSeg effectively encodes object motion, ensuring spatio-temporal consistency in segmentation results. To facilitate training and evaluation, we construct InstructVideo, a single- and multiple-object reasoning video segmentation benchmark. Comprehensive experiments demonstrate the effectiveness of the proposed components.
The Identity-conditioned Queries task is introduced, in which models are required to jointly associate and interpret an input video and a reference image of a person, and leverage this conditioning to address identity grounding, behavior understanding, and temporal reasoning, among other challenges.
Shibo Gao, Chongxiao Wang, Chenglong Huang et al.· 0 citations
STITCH is presented, a training-free method that divides a video into semantically meaningful temporal chunks that are computed once per video and reused across tasks, suggesting that reusable temporal abstraction is a promising direction for general video understanding.
Etienne Casanova, S. Brodjian, Pietro Perona· 0 citations
The Visually-Guided Disambiguation Aggregation Aggregation (VGD-Agg) framework is proposed, a framework based on a dual-branch fast-slow architecture that enhances discriminability via two learnable tokens and achieves state-of-the-art results on the proposed benchmarks.
Minghang Zheng, Jing Wei, Hong-Yi Yang et al.· 0 citations
It is demonstrated that prediction-level feedback substantially improves the reliability of training-free RVOS, with ReflexTrack, a training-free, feedback-driven agent that closes this loop at both spatial and temporal levels.
Yuanjia Li, Tianyang Xu, Tao Zhou et al.· arXiv.org· 0 citations
PhysMLLMs is a training-stage prior injection architecture that injects physics-inspired spatial continuity priors into Video MLLMs, demonstrating that the injected spatial prior improves video consistency without compromising image-level grounding or general multimodal capability.
Siyao Yan, Bo Han, Jisheng Dang et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.