This report summarizes the 8th Large-scale Video Object Segmentation (LSVOS) Challenge, held in conjunction with ECCV 2026. The challenge evaluates video segmentation in three complementary settings: complex semi-supervised video object segmentation on MOSEv2, text-guided referring video object segmentation on MeViSv2-Text, and audio-guided referring video object segmentation on MeViSv2-Audio. We describe the tasks and evaluation protocols and review the methods of the top three teams in each track. Across the nine leading solutions, foundation segmentation models are combined with target-aware memory, multimodal reasoning, explicit target-existence verification, agentic interaction, and corrective tracking. These systems illustrate a broader transition from single-model mask propagation toward modular pipelines that reason about object identity, query validity, and temporal reliability.
Chang Liu, Heng-Hui Ding, Ling-Yi Hong et al.· 0 citations
SAM2Dual is proposed, a training-free, plug-and-play inference-time enhancement that improves long-video robustness without updating model weights and Text-Aware Memory is presented, which extracts a compact word-level cue from early frames and uses text embeddings to reweight memory contributions based on semantic compatibility.
J. Kim, Changwon Lim· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.