Skip to content
Open access

A MULTIMODAL PIPELINE BRIDGING CAPTIONING AND OPEN-VOCABULARY DETECTION FOR ENHANCED VISIONLANGUAGE UNDERSTANDING

Jul 2026 · NLP & Big Data · 0 citations · 9 references

TL;DR

This paper presents a multimodal pipeline that combines vision-language captioning models and openvocabulary object detectors to investigate the impact of automatically generated textual prompts on semantic image understanding and reveals complementary behaviors between Grounding DINO and OWLv2.

Abstract

This paper presents a multimodal pipeline that combines vision-language captioning models and openvocabulary object detectors to investigate the impact of automatically generated textual prompts on semantic image understanding. The study evaluates several captioning models, including BLIP, BLIP-2, InstructBLIP, and LLaVA, in combination with two open-vocabulary detectors, OWLv2 and Grounding DINO. Experiments conducted on a representative subset of the COCO dataset show that prompt quality significantly influences detection performance and that post-processing operations, including label normalization and filtering, substantially improve semantic detection metrics. The results reveal complementary behaviors between Grounding DINO and OWLv2, highlighting the importance of prompt engineering and output refinement in multimodal vision-language pipelines. Rather than introducing a new detection architecture, this work provides a comparative analysis of the interactions between caption generation, prompt extraction, and open-vocabulary detection, offering insights for the design of future interactive vision-language systems.

Read PDF

Similar papers

Preprint Aug 2026

OPUS: A Simple yet Effective Unified Framework for Open-Vocabulary Detection

Recent unified open-vocabulary detection (OVD) supports heterogeneous prompts, including text queries, visual exemplars, and their combinations, but often rely on increasingly complex designs such as heavy cross-modal fusion, staged training, and iterative annotation pipelines. We revisit whether such complexity is necessary in the era of stronger foundation models. Our finding is that unified OVD can be made substantially simpler with semantic-rich visual representations and scalable grounding supervision. We present OPUS (\textbf{O}pen-vocabulary, \textbf{P}rompt-\textbf{U}nified, \textbf{S}imple), a unified detector supporting text, interactive visual, generic visual, and mixed prompting within one framework. OPUS adopts a simple three-part design. Its model architecture combines a semantic-rich visual encoder, built on a DINOv3-ConvNeXt-B backbone with efficient hybrid encoding, with a prompt-aware decoder that avoids prompt-specific branches for unified prompt reasoning. OPUS is trained with a one-stage text-visual training strategy with Instance-level Contrastive Alignment (ICA), and is supported by a SAM3-based single-pass data engine for heterogeneous grounding supervision. Experiments on COCO, LVIS-minival, and ODinW35 show that OPUS achieves state-of-the-art Visual-I performance, reaching 68.1/69.2/54.7 AP, while maintaining balanced Text and Visual-G accuracy. OPUS also turns mixed prompting from interference into complementarity, improving over text or visual prompt alone. These results show that simplicity and strong unified prompting capability can be achieved together.

Xiao-Yan Wei, Zhi-Min Yao, Rui-Lin Yang et al. · 0 citations
Review Open access Jul 2026

Multimodal Video Understanding: A Capability-Based Survey of Alignment, Expression, and Reasoning

A structured, comprehensive survey of the latest MVU progress is presented, establishing a novel three-tier taxonomy that categorizes existing studies into cross-modal alignment, multi-granularity semantic expression and multimodal reasoning.

Rongyong Zhao, Da Pu, Cuiling Li et al. · 0 citations
#computer vision Preprint Aug 2026

ReVA: A Region-Aware Visual Assistant for Visually Grounded Question Answering

Multimodal Large Language Models (MLLMs) have achieved remarkable progress in Visual Question Answering (VQA), yet they continue to struggle with questions requiring precise spatial reasoning and fine-grained visual understanding. These limitations often manifest as object, attribute, and spatial hallucinations, where models generate confident but visually unsupported responses due to insufficient region-level and fine-grained visual grounding. To address this challenge, we propose ReVA, a region-aware VQA model that employs a frozen CLIP ViT-L/14 Vision Transformer (ViT) and a Qwen2.5-7B-Instruct large language model (LLM) connected through a dual bridge that aligns both whole-image and region-level representations with the LLM's embedding space. The image bridge maps final transformer block features into image tokens. The region bridge maps cropped features from enriched intermediate features across ViT blocks so early texture and later object cues are more evident, into K region tokens for every bounding box. ReVA uses a detector stack that supplies automatic zero-shot bounding boxes that are both question-agnostic and question-dependent, using RAM++ (Recognize Anything Model), spaCy, and Grounding DINO. The image tokens and region tokens are concatenated as an LLM prompt prefix to jointly encode scene-level context and fine-grained regional evidence when answering questions. Evaluated on VQAv2, MMBench, POPE, and SEED-Bench, ReVA achieves 82.85% mean F1 on POPE, compared with 81.14% for an image-token baseline without region tokens. These results demonstrate that explicit region-aware visual representations reduce object hallucination and improve the factual grounding of MLLMs.

A. Senthil · 0 citations
Conference Jul 2026

Multimodal Context-Enriched Visual Representation Learning for Enhanced Vision–Language Image Captioning

Image captioning models can produce rapid Sentences, without visual relationships, or insert non-existing plausible objects. A common cause is to compress image evidence into visual symbols that carry a weak neighbourhood context. The multimodal context-enhanced visual representation learning framework (MCVRL) addresses this error mode by adding local neighbourhood descriptors, global scene tokens, prefix-conditioned visual doors and adaptive contextual corrections be-fore caption decoding. The encoder is trained with cross-entropy and contrast terms for image–text alignment. MSCOCO 2014’s Karpathy test classification results show that BLEU-4, METEOR, CIDEr and SPICE are more powerful captioning bases. The best configuration received a score of 1.352 of the CIDEr compared to 1.308 of BLIP-2 in the same evaluation protocol. The results of the ablation show that the most important contribution is the visual feature enriched by the context, followed by crossmodal gating and adaptive contextual attention. Qualitative examples show that objects with hallucinations are fewer and that spatial relationships are better recovered.

E. Divya, Johnson Kolluri, Kiran Siripuri · 0 citations
Open access Jul 2026

A Unified Multimodal Search Framework Using Generative AI and Image Understanding for Enhanced Information Retrieval

A unified multimodal search framework that integrates generative AI-based captioning and image understanding for improved retrieval, enabling more accurate, context-aware search in applications such as e-commerce, multimedia, and large-scale retrieval.

Saeed Alzahrani, Farah Mohammad, Nazar Hussain · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.