A New CLIP Based Domain-Specific Image-Text Alignment Approach for VQA and Image Captioning in Interpretation of Crowd Images
Abstract
With the development of attention mechanisms and transformer architectures, multi-mode problems such as automated image annotation and visual question-and-answer can be solved with high performance. The difficulty of these types of problems varies depending on the domain to which the image data belongs and the semantic focus of the text to be aligned. Recent studies have frequently focused on the natural language analysis of crowded images. The fact that existing pretrained multi-mode encoders such as CLIP are trained with contrastive training strategies based on the use of generalpurpose image-text alignment datasets makes it difficult to learn important semantic features in this field due to the complex structure of crowded images. In this study, a fine-tuning process of the CLIP encoder was performed using a powerful contrastive learning strategy with the visual question-andanswer dataset called COREVQA. As a result of this process, the core task of COREVQA, binary question-answer processing, was performed with an F-score of over 70% using a threshold-based method based on the similarity of text and visual features obtained from the CLIP encoder, achieving a performance superior to powerful multi-modal architectures such as GPT-4. Subsequently, to measure the contribution of the resulting domain-specific encoder to the crowd image captioning problem, synthetic captions of COREVQA images were obtained using the Gemini Flash API. These caption-image pairs were used to train a VLM architecture obtained through hybrid integration of CLIP and GPT-2. Ablation analysis of the developed captioning approach was performed for different VLM configurations using metrics such as ROUGE, BLEU, CIDEr, and METEOR. Comparative results show that our proposed domain-specific CLIP architecture has higher performance in image captioning and VQA problems in crowded images, according to all evaluation metrics.