A New CLIP Based Domain-Specific Image-Text Alignment Approach for VQA and Image Captioning in Interpretation of Crowd Images
With the development of attention mechanisms and transformer architectures, multi-mode problems such as automated image annotation and visual question-and-answer can be solved with high performance. The difficulty of these types of problems varies depending on the domain to which the image data belongs and the semantic...