SAM2Dual: Training-Free, Dual Memory for Long-Term Video Object Segmentation
SAM2Dual is proposed, a training-free, plug-and-play inference-time enhancement that improves long-video robustness without updating model weights and Text-Aware Memory is presented, which extracts a compact word-level cue from early frames and uses text embeddings to reweight memory contributions based on semantic compatibility.