Skip to content

Author

Prakhar Shukla

1 paper indexed here

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

Conference Jul 2026

Improving Trustworthiness in Visual Question Answering Via Question-Conditioned Cross-Modal Verification

Visual Question Answering (VQA) is a challenging cross-disciplinary task that combine natural language processing and computer vision together. VQA answer natural language questions based on images. Current VQA systems prone to generate plausible but incorrect responses without indicating uncertainty. To address this limitation, we propose a training-free cross-modal answer verification framework. This framework based on question-conditioned CLIP scoring as a post-hoc reliability estimator for vision-language models. The technique improves the ability to distinguish between right and wrong short-form VQA answers by evaluating candidate solutions jointly with the original question using type-aware distractor pools. Experiments on the VQAv2 validation set evaluate raw accuracy, verified accuracy, coverage, and accuracy gain under threshold-based selective prediction. Results shows consistent reliability improvements across diverse model architectures with gains ranging from +1.13 to +32.98 percentage points depending on coverage and model characteristics. The proposed framework is model-agnostic, requires no retraining or architectural modification, and improves the trustworthiness of multimodal systems by selectively filtering unreliable predictions. These findings show that question-conditioned CLIP verification provides an effective and scalable reliability layer for VQA systems.

Prakhar Shukla, Ankit Kumar, Pulkit Singh et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.