Skip to content

Evaluating LLM Robustness Under Domain-Specific Prompt Perturbations in Public Health Applications

Jul 2026 · arXiv.org · Vol abs/2607.06913 · 0 citations · 19 references
Computer Science

TL;DR

A domain specific robustness benchmark is proposed that evaluates LLMs under two perturbation types that commonly arise when non-clinical users interact with health AI systems: misinformation framing (MF) and layperson rewriting (LR).

Abstract

Large language models (LLMs) are increasingly applied in public health applications, yet their robustness to non-clinical user inputs remains underexplored. We propose a domain specific robustness benchmark that evaluates LLMs under two perturbation types that commonly arise when non-clinical users interact with health AI systems: misinformation framing (MF), where prompt might be injected by false health claims, and layperson rewriting (LR), where patients describe symptoms in everyday language rather than medical terminology. Our goal is to evaluate the stability of LLMs under these perturbation. Experiments show that MF degrades accuracy by 7.2 pp on average with prediction flip rates of 9-38 percent, even when claims are explicitly labelled as unsupported; LR causes only 1.4 pp degradation. These findings highlight two distinct deployment risks in public health settings: models may produce incorrect outputs when users unintentionally carry misinformation into their queries, and may misinterpret clinically relevant details when patients use informal language. Both risks call for perturbation-aware robustness evaluation beyond clean baseline benchmark

View source

Similar papers

Conference Open access Jul 2026

When Confidence Fails: Overconfidence in LLMS Under Uncertainty and Missing Clinical Information

An evaluation framework based on the MedMCQA dataset consisting of two complementary uncertainty settings is proposed, which introduces linguistic uncertainty cues through prompt modifications to simulate ambiguous clinical contexts and observes significant variation across models in their ability to abstain when the correct answer is unavailable.

Maryam Tahermazandarani, Adnan Mahmood, Fahmida Islam et al. · 0 citations
Open access Jul 2026

Addressing benchmarking gaps in large language models for health and medicine with dynamic red-teaming

A Dynamic, Automatic and Systematic red-teaming audit framework that continuously stress-tests LLMs for health across four safety-critical axes: robustness, privacy, bias and hallucination, which provides a scalable framework for surfacing latent risks before such systems are deployed in consumer-facing health assistants and broader clinical workflows.

Jiazhen Pan, Bailiang Jian, Paul Hager et al. · 0 citations
Book Open access Feb 2026

LiveMedBench: A Contamination-Limited Medical Benchmark for LLMs with Automated Rubric Evaluation

A Multi-Agent Clinical Curation Framework that filters raw data noise and validates clinical integrity against evidence-based medical principles and develops an Automated Rubric-based Evaluation Framework that decomposes physician responses into granular, case-specific criteria, achieving substantially stronger alignment with expert physicians than LLM-as-a-Judge.

Zhiling Yan, D. Song, Zhen Fang et al. · 8 citations · ⚡2
#small language model Preprint Aug 2026

Future Querying: Can LLMs Serve as Implicit Medical World Models?

This work introduces future querying, a paradigm that probes whether large language models can function as implicit medical world models by evaluating their ability to answer time-indexed clinical queries about a patient's future, and shows that small, locally fine-tuned open-weight models can match or approach larger proprietary systems, making the framework suitable for privacy-preserving, on-premise deployment.

Siri Willems, James Butterworth, L. Goetschalckx et al. · 0 citations
Jul 2026

MamaBench: Benchmarking LLM Robustness in Maternal and Child Health Diagnosis through Counterfactual Clinical Perturbation

Evidence-Anchored RAG is proposed (EA-RAG), a three-stage retrieval method that replaces aggregate similarity with an evidence coverage objective through clinical parameter extraction, coverage auditing, and contrastive sub-queries, and confirms that counterfactual robustness in clinical AI remains an open challenge.

Thanni Adewuyi, Anuoluwa Sotome, Samuel Okoko et al. · 0 citations
Open access Jul 2026

Mitigating medical bias in large language models by prompt engineering: an empirical study of effectiveness and trade-offs.

Five widely used prompting strategies across five influential LLMs in the latest medical bias benchmark reveal substantial heterogeneity in both effectiveness and overhead across models, with no strategy proving universally effective and some even exacerbating bias.

Ying Xiao, Zhenpeng Chen, Jie M. Zhang · 1 citation

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.