Skip to content
Preprint

ViCo-SAM3: Vision-Conditioned Alignment for Open-Vocabulary Camouflaged Object Segmentation

Sep 2026 · 1 citation · 21 references
Computer Science

TL;DR

This work proposes ViCo-SAM3, a Vision-Conditioned alignment framework designed for OVCOS, which introduces vision-conditioned (ViCo) module, which dynamically modulates text embeddings with global visual context, enabling textual representations to adapt to the current image content and thereby effectively bridges the semantic gap between vision and text.

Abstract

Open-vocabulary camouflaged object segmentation (OVCOS) aims to segment unseen camouflaged objects under text guidance. We observe that SAM3 still suffers from a pronounced semantic gap between global textual semantics and fine-grained pixel-level visual cues in OVCOS. Meanwhile, fully fine-tuning the text encoder introduces heavy parameter overhead and risks overfitting to training categories, which compromises open-vocabulary representation flexibility. To address these issues, we propose ViCo-SAM3, a Vision-Conditioned alignment framework designed for OVCOS. Specifically, we introduce vision-conditioned (ViCo) module, which dynamically modulates text embeddings with global visual context, enabling textual representations to adapt to the current image content and thereby effectively bridging the semantic gap between vision and text. Building on this, we further design a vision-conditioned cross-modal binding (ViCoBind) module to enhance cross-modal interaction and semantic alignment between visual and textual representations. Without bells and whistles, ViCo-SAM3 achieves state-of-the-art performance on the OVCamo benchmark and demonstrates strong generalization.

View source

Similar papers

Open access Sep 2026

MODdapter: Spatially-aware text embeddings for zero-shot semantic segmentation.

Vision-language models show promise in zero-shot semantic segmentation, but a key challenge is the disconnect between text and visual features. While text embeddings can roughly localize unseen objects, they often lack the fine-grained detail necessary for accurate segmentation, leading to oversegmentation or undersegm...

Jia-Xiang Fang, Shi-Qiang Ma, Jing Wang et al. · 0 citations
Conference Open access Sep 2026

Open-Vocabulary Object 6D Pose Estimation via Modulated Textual Semantics

Estimating the 6D pose of novel objects without CAD models or video sequences remains a challenging problem. Recent works explore text-driven approaches to address this challenge in an open-vocabulary manner. However, these methods typically treat text embeddings as static priors, which lack the flexibility to adapt to...

Zixuan Sun, H. Shuai, Qing-Shan Liu · 0 citations
Conference Open access Sep 2026

Auxiliary text-guided image restoration for image-text matching

Image-Text Matching (ITM) aims to establish deep semantic associations between visual content and textual descriptions. Existing methods usually have discrimination issues because of fine-grained semantic deviations, so it's hard to capture the complex correspondences between cross-modal entries. Only relying on alignm...

Kuang-Rong Hao · 0 citations
Sep 2026

XMask3D++: Cross-Modal Mask Reasoning for Open Vocabulary 3D Segmentation.

Existing methodologies in open vocabulary 3D semantic segmentation primarily concentrate on establishing a unified feature space encompassing 3D, 2D, and textual modalities. Nevertheless, traditional techniques such as global feature alignment or vision-language model distillation tend to impose only approximate corres...

Ziyi Wang, Yan-Bo Wang, Xu-Min Yu et al. · 0 citations
Preprint Aug 2026

LAD-COD: Language-Aligned Dense Perception for Camouflaged Object Detection

Camouflaged object detection (COD) aims to segment objects that exhibit high visual similarity to their surroundings, which reduces foreground-background discriminability and weakens boundary evidence across appearance, texture, and structure. Such limitations motivate the use of instruction-conditioned semantics as to...

Shangye Song, Tianzhi Zhu, Syed Ariff Syed Hesham et al. · 0 citations
Open access Aug 2026

Learning from multimodal pseudo-labels for robust open-vocabulary instance and panoptic segmentation

This work addresses the challenge of open-vocabulary instance segmentation (OVIS) and open-set panoptic segmentation (OSPS), which aim to recognize both predefined and unseen object categories without exhaustive human annotations. Existing methods often suffer from noisy pseudo-masks, limited visual-textual grounding,...

Duy Tran Thanh, Yee-Jin Lee, Byeongkeun Kang · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.