TNG-CLIP:Training-Time Negation Data Generation for Negation Awareness of CLIP
Multimodal interactive systems increasingly rely on vision-language models (VLMs), such as CLIP, to mediate between users and visual content. These VLMs power image search and retrieval interfaces where users specify what they want in natural language. However, such systems suffer from a basic aspect of human communica...