Jul 2026
Fine-grained CLIP fine-tuning with self-annotated region alignment
SFF-CLIP (Self-annotated Fine-grained Fine-tuning for CLIP), which only uses image-text pairs as input to boost the fine-grained representation ability in the CLIP fine-tuning, while maintaining the global visual-semantic consistency.
Chen-Yang Zhao, Wei Lin, Antoni B. Chan et al.
· arXiv.org · 0 citations