Skip to content
Conference

Improving Trustworthiness in Visual Question Answering Via Question-Conditioned Cross-Modal Verification

Jul 2026 · 2026 4th International Conference on Sustainable Computing and Smart Systems (ICSCSS) · pp. 2094-2100 · 0 citations · 17 references

Abstract

Visual Question Answering (VQA) is a challenging cross-disciplinary task that combine natural language processing and computer vision together. VQA answer natural language questions based on images. Current VQA systems prone to generate plausible but incorrect responses without indicating uncertainty. To address this limitation, we propose a training-free cross-modal answer verification framework. This framework based on question-conditioned CLIP scoring as a post-hoc reliability estimator for vision-language models. The technique improves the ability to distinguish between right and wrong short-form VQA answers by evaluating candidate solutions jointly with the original question using type-aware distractor pools. Experiments on the VQAv2 validation set evaluate raw accuracy, verified accuracy, coverage, and accuracy gain under threshold-based selective prediction. Results shows consistent reliability improvements across diverse model architectures with gains ranging from +1.13 to +32.98 percentage points depending on coverage and model characteristics. The proposed framework is model-agnostic, requires no retraining or architectural modification, and improves the trustworthiness of multimodal systems by selectively filtering unreliable predictions. These findings show that question-conditioned CLIP verification provides an effective and scalable reliability layer for VQA systems.

View source

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.