A domain specific robustness benchmark is proposed that evaluates LLMs under two perturbation types that commonly arise when non-clinical users interact with health AI systems: misinformation framing (MF) and layperson rewriting (LR).
Abstract
Large language models (LLMs) are increasingly applied in public health applications, yet their robustness to non-clinical user inputs remains underexplored. We propose a domain specific robustness benchmark that evaluates LLMs under two perturbation types that commonly arise when non-clinical users interact with health AI systems: misinformation framing (MF), where prompt might be injected by false health claims, and layperson rewriting (LR), where patients describe symptoms in everyday language rather than medical terminology. Our goal is to evaluate the stability of LLMs under these perturbation. Experiments show that MF degrades accuracy by 7.2 pp on average with prediction flip rates of 9-38 percent, even when claims are explicitly labelled as unsupported; LR causes only 1.4 pp degradation. These findings highlight two distinct deployment risks in public health settings: models may produce incorrect outputs when users unintentionally carry misinformation into their queries, and may misinterpret clinically relevant details when patients use informal language. Both risks call for perturbation-aware robustness evaluation beyond clean baseline benchmark
An evaluation framework based on the MedMCQA dataset consisting of two complementary uncertainty settings is proposed, which introduces linguistic uncertainty cues through prompt modifications to simulate ambiguous clinical contexts and observes significant variation across models in their ability to abstain when the correct answer is unavailable.
Maryam Tahermazandarani, Adnan Mahmood, Fahmida Islam et al.· International Conference on...· 0 citations
A Dynamic, Automatic and Systematic red-teaming audit framework that continuously stress-tests LLMs for health across four safety-critical axes: robustness, privacy, bias and hallucination, which provides a scalable framework for surfacing latent risks before such systems are deployed in consumer-facing health assistants and broader clinical workflows.
Jiazhen Pan, Bailiang Jian, Paul Hager et al.· Nature Health· 0 citations
A Multi-Agent Clinical Curation Framework that filters raw data noise and validates clinical integrity against evidence-based medical principles and develops an Automated Rubric-based Evaluation Framework that decomposes physician responses into granular, case-specific criteria, achieving substantially stronger alignment with expert physicians than LLM-as-a-Judge.
Zhiling Yan, D. Song, Zhen Fang et al.· Proceedings of the 32nd ACM...· 8 citations· ⚡2
This work introduces future querying, a paradigm that probes whether large language models can function as implicit medical world models by evaluating their ability to answer time-indexed clinical queries about a patient's future, and shows that small, locally fine-tuned open-weight models can match or approach larger proprietary systems, making the framework suitable for privacy-preserving, on-premise deployment.
Siri Willems, James Butterworth, L. Goetschalckx et al.· 0 citations
Evidence-Anchored RAG is proposed (EA-RAG), a three-stage retrieval method that replaces aggregate similarity with an evidence coverage objective through clinical parameter extraction, coverage auditing, and contrastive sub-queries, and confirms that counterfactual robustness in clinical AI remains an open challenge.
Thanni Adewuyi, Anuoluwa Sotome, Samuel Okoko et al.· arXiv.org· 0 citations
Five widely used prompting strategies across five influential LLMs in the latest medical bias benchmark reveal substantial heterogeneity in both effectiveness and overhead across models, with no strategy proving universally effective and some even exacerbating bias.
Ying Xiao, Zhenpeng Chen, Jie M. Zhang· Philosophical transactions....· 1 citation
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.