A MULTIMODAL PIPELINE BRIDGING CAPTIONING AND OPEN-VOCABULARY DETECTION FOR ENHANCED VISIONLANGUAGE UNDERSTANDING
This paper presents a multimodal pipeline that combines vision-language captioning models and openvocabulary object detectors to investigate the impact of automatically generated textual prompts on semantic image understanding and reveals complementary behaviors between Grounding DINO and OWLv2.