ViCo-SAM3: Vision-Conditioned Alignment for Open-Vocabulary Camouflaged Object Segmentation
This work proposes ViCo-SAM3, a Vision-Conditioned alignment framework designed for OVCOS, which introduces vision-conditioned (ViCo) module, which dynamically modulates text embeddings with global visual context, enabling textual representations to adapt to the current image content and thereby effectively bridges the...