Skip to content
Open access

Accuracy Overstates Evidence Grounding and Abstention Reliability in Mammography Vision-Language Models

Sep 2026 · medRxiv · 0 citations
Medicine

TL;DR

An evidence-grounded selective evaluation benchmark that evaluates Pathology classification and Abnormality identification together with label-aware lesion localization suggests that answer accuracy and abstention behavior can substantially overstate the reliability of current VLM when predictions are not verified against the visual evidence on which they should depend.

Read PDF

Similar papers

#machine learning Preprint Sep 2026

Contamination, Prior, or Evidence? Decomposing and Training Evidence Use in Whole-Slide Vision-Language Models

Pathology vision-language models (VLMs) are conventionally evaluated by accuracy, but accuracy alone does not measure evidence use: it may conflate dataset contamination, prior knowledge, and image evidence. In a motivating study of lymph-node metastasis prediction, we found that most public pathology VLMs showed minim...

Wen-Hao Zhang, Zhong-Liang Zhou, Shi-Yuan Zhang et al. · 0 citations
Open access Sep 2026

An Automated, Contamination-Controlled VQA Benchmark for Evaluating Vision-Language Models on 3D Oncology Imaging

Vision-language models (VLMs) are increasingly applied to medical imaging, yet public benchmarks may reward memorization over perception: their images and questions can enter pretraining corpora, and many items remain answerable from question text alone. We present an automated, agent-driven pipeline that builds multip...

Bo Liu, Han Gu, Xiang-Rui Li et al. · 0 citations
Aug 2026

Do Accuracy Gains Reflect Genuine Visual Understanding? A Multi-Model Evaluation of Vision-Language Models in Thoracic Imaging.

Even at the highest accuracy observed here, answer-level performance can overstate visual understanding, pointing to the need for more appropriate evaluation methods that assess whether a model's reasoning is faithful to the image, not only whether its answer is correct.

Jiyoung Song, Won Gi Jeong, Hongseok Ko et al. · 0 citations
Open access Aug 2026

CBEC: a simple retrieval-based framework for population-grounded prediction reliability estimation in chest X-ray diagnosis

It is suggested that case-level consistency provides a meaningful practical signal for reliability estimation in medical imaging, whose relationship to the classical epistemic–aleatoric decomposition warrants further theoretical investigation across broader clinical settings.

Muhannad Faleh Alanazi, B. Z. Shakhreet, Hattan Ali A. Asiri et al. · 0 citations
#artificial intelligence Preprint Sep 2026

Knowing When Not to Answer: Abstention and Refusal Reasoning in Vision--Language Models

Many medical conditions require diagnosis through detailed, multi-context clinical assessment rather than from visual appearance alone. Despite this, vision-language models (VLMs) are increasingly queried to interpret images in ways that touch on medical or diagnostic judgments, raising safety concerns when such infere...

Karan Dua, Amit Agarwal, Hitesh Laxmichand Patel et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.