Skip to content
Open access

Toward Reasoning-Centric Video Object Segmentation via Multi-Modal Large Language Models

Jul 2026 · IEEE Transactions on Image Processing · Vol 35, pp. 8206-8218 · 0 citations · 73 references
Medicine Computer Science

Abstract

Referring Video Object Segmentation (RVOS) aims to segment the target objects specified in human instructions. Previous approaches typically rely on explicit human instructions that contain target categories or salient appearance descriptions. These approaches tend to fail when the instructions require temporal video understanding and complex relational reasoning. In this work, we present RViSeg, a reasoning-centric video object segmentation model that leverages the reasoning capability of Multi-modal Large Language Models (MLLM) to handle complex queries. The primary challenge lies in enabling MLLM to perform efficient pixel-level video perception. To tackle this challenge, we introduce a novel Spatial Token Merge (STM) module that consolidates lengthy video tokens into compact region-level clusters, while preserving essential spatial details. This structured representation enables MLLM to infer user intention by interleaving spatial and temporal visual information. Furthermore, we propose a Query-based Target Retrieval (QTR) module that utilizes learnable tokens as the target identity for mask prediction. By propagating these instance-specific tokens both intra-clip and inter-clip, our RViSeg effectively encodes object motion, ensuring spatio-temporal consistency in segmentation results. To facilitate training and evaluation, we construct InstructVideo, a single- and multiple-object reasoning video segmentation benchmark. Comprehensive experiments demonstrate the effectiveness of the proposed components.

Read PDF

Similar papers

Review Aug 2026

I Seek You in Videos: Identity-Conditioned Queries for Person-Centric Video Reasoning

The Identity-conditioned Queries task is introduced, in which models are required to jointly associate and interpret an input video and a reference image of a person, and leverage this conditioning to address identity grounding, behavior understanding, and temporal reasoning, among other challenges.

Shibo Gao, Chongxiao Wang, Chenglong Huang et al. · 0 citations
#small language model Preprint Aug 2026

Training-Free Temporal Abstraction for General Video Understanding

STITCH is presented, a training-free method that divides a video into semantically meaningful temporal chunks that are computed once per video and reused across tasks, suggesting that reusable temporal abstraction is a promising direction for general video understanding.

Etienne Casanova, S. Brodjian, Pietro Perona · 0 citations
Preprint Aug 2026

ID-VTG: Image-Disambiguated Video Temporal Grounding

The Visually-Guided Disambiguation Aggregation Aggregation (VGD-Agg) framework is proposed, a framework based on a dual-branch fast-slow architecture that enhances discriminability via two learnable tokens and achieves state-of-the-art results on the proposed benchmarks.

Minghang Zheng, Jing Wei, Hong-Yi Yang et al. · 0 citations
Jul 2026

ReflexTrack: A Feedback-Driven Agent for Training-Free Referring Video Object Segmentation

It is demonstrated that prediction-level feedback substantially improves the reliability of training-free RVOS, with ReflexTrack, a training-free, feedback-driven agent that closes this loop at both spatial and temporal levels.

Yuanjia Li, Tianyang Xu, Tao Zhou et al. · 0 citations
#small language model Preprint Aug 2026

PhysMLLMs: Spatial Priors for Unified Referring Segmentation and Grounded Reasoning of Images and Videos

PhysMLLMs is a training-stage prior injection architecture that injects physics-inspired spatial continuity priors into Video MLLMs, demonstrating that the injected spatial prior improves video consistency without compromising image-level grounding or general multimodal capability.

Siyao Yan, Bo Han, Jisheng Dang et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.