Clinical warrant in medical LLM evaluation
Abstract
Medical large language models are evaluated through factuality, physician preference, source support, retrieval quality, safety, and clinical utility. These measures describe content and performance while leaving unresolved whether a generated claim meets the requirements for a specified clinical use. This article proposes clinical warrant as a use-specific judgment for a particular output–use pair. Clinical use includes reliance within a patient-related encounter, task, or care pathway, including education that can shape understanding or later action. The proposed use and the standards used to judge it are documented separately. For action-guiding outputs, clinical appropriateness is a gateway. Five further conditions examine role, case information, evidence, preserved limits, and authority/accountability. The framework asks whether these requirements are met at the point of proposed reliance, records failed or unresolved conditions and the resulting disposition, and tests whether a relevant change stops or redirects use. Fifteen constructed examples across five domains distinguish established, not established, and indeterminate warrant.