Skip to content
Preprint

Evaluating the Diagnostic Robustness of Vision-Language Models Under Visual and Textual Perturbations

Aug 2026 · 0 citations · 17 references
Computer Science

TL;DR

This work utilizes a histopathology-validated brain MRI dataset to systematically assess the diagnostic robustness of four VLM families under evidence-preserving perturbations, and reveals significant vulnerabilities in presentation-order stability.

Abstract

Standard accuracy metrics for VLMs often mask significant reliability failures in sensitive domains. In this work, we utilize a histopathology-validated brain MRI dataset to systematically assess the diagnostic robustness of four VLM families under evidence-preserving perturbations. By reordering anatomical slices and swapping target label positions, we evaluate whether models maintain consistent predictions when clinical evidence remains invariant. Our results reveal significant vulnerabilities in presentation-order stability, with models exhibiting prediction flips in up to 48.9% of cases under simple sequence reversals. We further identify a textual selection bias, where label reordering triggers inconsistent diagnoses in up to 67.8% of cases despite identical visual inputs. Negative-control tests further reveal diagnostic overcommitment: models generate categorical diagnoses in up to 76.1% of cases after expert-annotated lesion slices are removed. These results demonstrate that high accuracy can overestimate clinical reliability, masking sensitivity to sequential presentation and textual framing that is not captured by aggregate accuracy. Our findings highlight the necessity of stability-based metrics for the deployment of VLMs in safety-critical clinical applications. Our evaluation data and code will be made public upon acceptance.

View source

Similar papers

Open access Sep 2026

Accuracy Overstates Evidence Grounding and Abstention Reliability in Mammography Vision-Language Models

An evidence-grounded selective evaluation benchmark that evaluates Pathology classification and Abnormality identification together with label-aware lesion localization suggests that answer accuracy and abstention behavior can substantially overstate the reliability of current VLM when predictions are not verified agai...

B. Qu, W. Liu, M. Murrow et al. · 0 citations
Preprint Oct 2026

PathLang: A Language-Centered Benchmark for Vision-Language Models in Computational Pathology

Pathology vision-language models (VLMs) have shown strong visual perception ability, but their robustness in the language domain remains poorly characterized. Existing pathology VLM benchmarks largely rely on canonical closed-set prompts or perturb only generic templates, treating language as a fixed evaluation compone...

Fang-Qi Cheng, Kuo Gong, Shan Liu et al. · 0 citations
#artificial intelligence Preprint Sep 2026

Beneath the Scores: Rethinking Hallucination Evaluation for Video Understanding Models

Video understanding is increasingly performed by multi-stage LLM agents that separate temporal grounding, visual observation, and reasoning. Yet these stages are typically evaluated on different benchmarks and distributions, making it difficult to determine where hallucinations originate. We first organize existing ben...

Shu-Zhi Gong, F. Sun, Yuansan Liu · 0 citations
Open access Sep 2026

Martingale Doppelgänger-Eval: A Specification-Driven Interventional Audit Algorithm for Visual Evidence Use in Vision–Language Models

Assessing chart understanding requires distinguishing responses to local visual evidence from associations with a chart’s overall shape. We present Martingale Doppelgänger-Eval, an interventional audit framework that connects executable edit specifications to benchmark generation, validity checks and response estimatio...

Zi-Yao Wang, Svetlozar T. Rachev · 0 citations
#artificial intelligence Review Sep 2026

See, Measure, and Reason: Learning Visually Grounded Reasoning in Pathology

Pathological assessment relies on recognizing fine-grained visual details in histological images. Vision-language models (VLMs) increasingly support pathology interpretation, yet their ability to perceive these details remains inadequate. This weakness leads to inaccurate cellular observations that can persist even whe...

Cheng-Yang Zhang, Wen-Chuan Zhang, Bo Li et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.