Claude Fable 5, Anthropic's most capable publicly available model, is evaluated across eight biomedical benchmarks, four text and four multimodal, using deterministic scoring against fixed answer keys throughout.
Abstract
Frontier language models are increasingly evaluated on biomedical benchmarks, but two problems undermine most published evaluations: legacy benchmarks are near-saturated, and open-ended responses are graded by other language models. We evaluate Claude Fable 5, Anthropic's most capable publicly available model, across eight biomedical benchmarks, four text and four multimodal, using deterministic scoring against fixed answer keys throughout. We include two Claude predecessors and GPT-5 as baselines. Refusal is tracked as a distinct outcome in every result table. That decision produces the paper's central finding. Fable 5 refuses between 8.0% and 99.4% of questions depending on the benchmark, a pattern absent in both predecessors and in GPT-5. Once refused items are excluded from the denominator, Fable 5's accuracy exceeds or meets every other model on every benchmark in this study. We identify two distinguishable refusal patterns: one concentrating in basic-science and mechanism content across MedQA and MedXpertQA MM, confirmed independently on two benchmarks using each benchmark's own category labels; and a separate disease-domain pattern on RareBench, where inborn metabolic disease presentations are refused near-universally while adult-onset autoimmune presentations are not. The primary constraint on Fable 5's biomedical usefulness is willingness to engage, not capability once it does.
Multiple-choice medical benchmarks are increasingly saturated, and recent rubric-based evaluations such as HealthBench have shown that open-ended clinical performance is far from solved - its"Hard"subset top score remains 32%. We present a small, deliberately difficult evaluation dataset of five clinician-authored clinical scenarios spanning four specialties (anaesthesia, internal/family medicine, emergency medicine, and obstetrics), each accompanied by an atomic, weighted, MECE rubric (25-62 criteria per task; 184 criteria total) authored from a clinician-drafted golden answer. We evaluate three frontier models: GPT 5.4, Claude Opus 4.7, and Gemini 3.1 Pro. Mean rubric pass rates were 0.47 (Claude), 0.39 (GPT), and 0.37 (Gemini). The central finding is an inversion of clinical priority: the highest-weighted (weight-5, critical) criteria passed at only 32.4-41.7%, while low-stakes weight-1 criteria passed at 80-90%. 56 of 108 critical (weight-5) criteria (52%) were satisfied by no model. Three LLM autoraters reproduced expert met/not-met labels on 92.8-94.7% of 552 graded criteria. We position this as a methods-and-preliminary-findings contribution: the five tasks demonstrate a scalable, defensible pipeline ready to develop into a large-scale benchmark.
Samiha A. Ismail, Fan X. Chen, Ali Merali· 0 citations
VITA's advantages in accuracy and completeness persisted under the neutral judge; its communication scores were lower, and this results indicate that a purpose-built clinical RAG system remains competitive with frontier LLMs on an open benchmark, consistent with corpus specificity as a design variable that improves grounding at some cost to communication polish.
Praveen Reddy, C. Mandke, Suvrankar Datta et al.· 0 citations
This entry-level tutorial aims to equip healthcare professionals with the tools necessary to effectively integrate LLMs into clinical practice, ensuring that these powerful technologies are applied in a safe, reliable, and impactful manner.
Qiao Jin, Nicholas Wan, Robert Leaman et al.· Nature Protocols· 1 citation
Across medical benchmarks, MedLLM shows a pattern visible only at sub-billion scale: medical competence does not degrade uniformly under compression but splits by task type and dissociation is masked at 7B, where both capabilities are present, and surfaces only when capacity is scarce.
M. R. Rahman, Asim Ahmed, Mihan Mohagheghzadeh et al.· 0 citations
Clinical LLMs carry an ordered evidence-strength signal they do not express, so their stated grades fail to convey a claim's support even when it is recoverable from their representations and text.
Background Large language models (LLMs) perform well on medical examinations, but they are almost always evaluated as test-takers, judged on the accuracy of their answers. Far less is known about their reliability as assessment-automation tools, for instance in extracting examination metadata, or about whether single-run, accuracy-based benchmarks can adequately characterize such tools. Methods Twelve mainstream LLMs (8 domestic, 4 international) were evaluated on extracting eight metadata fields from a 60-item endocrinology examination, each performing the identical task in three independent runs under realistic web-based conditions. Performance was separated into completion rate, conditional accuracy, and an all-fields-correct task success rate (TSR), with cross-run variance treated as a primary outcome. Conditional accuracy and TSR were compared within models, and domestic versus international models were compared on TSR. Results Conditional accuracy was uniformly high, yet TSR was sharply bimodal: five models scored 90% or above and seven below 70%, with none in between. Reliability, not mean accuracy, was decisive. Several models with perfect conditional accuracy collapsed when one or two of their three runs failed entirely, a pattern that single-run evaluation would miss. Failures spanned both model behaviour and platform-level constraints, the latter not removable by prompting. Model origin did not predict performance. On one item, four models independently changed the examiner’s assigned cognitive level to one of their own, a “silent relabeling” with direct implications for item-bank integrity. Conclusion For automated assessment metadata extraction, reliability rather than accuracy determines whether an LLM can be deployed, and national origin is not a meaningful predictor. Such tools should be evaluated by repeated runs reporting cross-run variance, and adopted through task-specific piloting within a workflow that reserves human judgement for semantic and normative fields.
Hui Zhang, Lihui Qu, Hongbo Bai et al.· Frontiers in Artificial Inte...· 0 citations