Skip to content

Author

Tanupriya Choudhury

2 papers indexed here

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

Conference Jul 2026

Improving Trustworthiness in Visual Question Answering Via Question-Conditioned Cross-Modal Verification

Visual Question Answering (VQA) is a challenging cross-disciplinary task that combine natural language processing and computer vision together. VQA answer natural language questions based on images. Current VQA systems prone to generate plausible but incorrect responses without indicating uncertainty. To address this limitation, we propose a training-free cross-modal answer verification framework. This framework based on question-conditioned CLIP scoring as a post-hoc reliability estimator for vision-language models. The technique improves the ability to distinguish between right and wrong short-form VQA answers by evaluating candidate solutions jointly with the original question using type-aware distractor pools. Experiments on the VQAv2 validation set evaluate raw accuracy, verified accuracy, coverage, and accuracy gain under threshold-based selective prediction. Results shows consistent reliability improvements across diverse model architectures with gains ranging from +1.13 to +32.98 percentage points depending on coverage and model characteristics. The proposed framework is model-agnostic, requires no retraining or architectural modification, and improves the trustworthiness of multimodal systems by selectively filtering unreliable predictions. These findings show that question-conditioned CLIP verification provides an effective and scalable reliability layer for VQA systems.

Prakhar Shukla, Ankit Kumar, Pulkit Singh et al. · 0 citations
Conference Jul 2026

AE-CCAF: Cross-Attention Guided Ensemble Learning for Robust Multi-Class Skin Lesion Classification

Proper and early classification of skin lesions is crucial for the effective diagnosis of melanoma, but remains challenging due to class imbalance, class similarity, and artefacts in dermoscopic images. This paper proposes AE-CCAF, an adaptive ensemble framework that combines heterogeneous deep backbones with a cross-attention fusion mechanism to robustly classify multi-class skin lesions. The model uses complementary representations from EfficientNet, Swin Transformer, and ConvNeXt, along with a cross-attention module to dynamically weight inter-feature interactions. As a measure to tackle data imbalance, an asymmetric focal loss using class-sensitive weighting is added. Extensive experiments on the ISIC 2018 benchmark indicate that AE-CCAF achieves a macro-AUC of 0.9487 and a balanced accuracy of 83.7%, outperforming current state-of-the-art algorithms. The proposed solution enhances the sensitivity for clinically urgent cases of melanoma while maintaining very high specificity across all levels. These findings point to the success of attention-guided ensemble learning to provide reliable and scalable dermatological diagnosis.

Pawan Kumar, Tanupriya Choudhury, Roohi Sille et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.