Author

Gianluca Moro

1 paper indexed here

Fetches their full publication history.

Not the right person? Other researchers publish under this name.

Conference Open access 2026

LLMs (Almost) Never Abstain Under Medical Uncertainty

Medical multiple-choice question answering (MCQA) benchmarks implicitly assume that large language models (LLMs) should always commit to an answer. However, in clinical practice, uncertainty is pervasive and abstaining is often the safest action. We introduce MedQAbstain , a benchmark explicitly designed to evaluate medical abstention under uncertainty. MedQAbstain repurposes standard medical MCQA datasets by removing the gold answer and introducing an explicit “I ab-stain” option, framed as a safety-critical decision with clinical consequences. The benchmark supports systematic analysis across ab-stention regimes, distractor complexity, and input modalities, and elicits self-reported model confidence to study calibration. Across all settings, we find that state-of-the-art LLMs systematically overcommit, rarely abstaining even when the question itself is hidden. These results reveal a fundamental mismatch between LLM behavior and clinical norms, highlighting ab-stention as a critical but overlooked dimension of medical decision-making evaluation. 1

Alessio Cocchieri, Luca Ragazzi, Giuseppe Tagliavini et al. · 2 citations