2026· Proceedings of machine learning research· Vol 297, pp. 1023 - 1046· 0 citations
Medicine
TL;DR
A playbook for fine-tuning LLMs on de-identified clinical notes from patients with pancreatic cancer, spanning both pre-diagnosis and on-treatment settings is presented, and the use of machine-generated annotations to augment limited expert labels is examined, showing that balanced mixtures of synthetic and human data can enhance fine-tuned models.
Abstract
Large language models (LLMs) are increasingly applied to clinical notes, but guidance on how to adapt open-source models to specific tasks and manage annotation quality at scale is limited. We present a playbook for fine-tuning LLMs on de-identified clinical notes from patients with pancreatic cancer, spanning both pre-diagnosis and on-treatment settings. We evaluate prompting strategies, contrast open-source models with GPT-4o, and explore disease-level versus task-specific adaptation. A key contribution is an LLM-assisted adjudication workflow in which models flag notes where predictions consistently conflict with initial human labels. This approach concentrated expert review on a small fraction of cases while identifying many true annotation errors, ultimately improving downstream model performance. We further examine the use of machine-generated annotations to augment limited expert labels, showing that balanced mixtures of synthetic and human data can enhance fine-tuned models. Our findings provide practical guidance for deploying open-source LLMs in clinical contexts, offering strategies to improve accuracy, reduce annotation burden, and enable privacy-preserving, site-adapted clinical natural language processing (NLP).
A Multi-Agent Clinical Curation Framework that filters raw data noise and validates clinical integrity against evidence-based medical principles and develops an Automated Rubric-based Evaluation Framework that decomposes physician responses into granular, case-specific criteria, achieving substantially stronger alignment with expert physicians than LLM-as-a-Judge.
Zhiling Yan, D. Song, Zhen Fang et al.· Proceedings of the 32nd ACM...· 8 citations· ⚡2
Early screening of chronic kidney disease (CKD) is essential for preventing irreversible progression; however, many machine learning (ML)-based screening methods remain difficult to deploy in community and resource-limited screening settings due to their reliance on large labeled datasets, resource-intensive pathology tests, or high-dimensional clinical features, and limited robustness to population and distributional shifts. This study examines the feasibility of using large language models (LLMs) for early-stage CKD screening in a zero-shot setting, without dataset-specific training. We propose a feature-guided zero-shot framework that evaluates LLM performance using a selected set of clinically meaningful, readily available community-based features, rather than exhaustive clinical inputs. Feature selection was guided by ML-based analysis to identify a compact, clinically relevant subset of variables. Tabular patient records were subsequently serialized into text using standardized prompt templates to enable zero-shot inference. The zero-shot performance of four LLMs (LLaMA-3, Qwen-3, Mistral, and GPT-4o-mini) was evaluated using both the full feature set and the selected subset. Generalizability was assessed across three heterogeneous CKD datasets spanning three countries. Across models and datasets, the selected feature set yielded consistent and statistically significant improvements in balanced accuracy and probability estimates, achieving performance levels suitable for screening purposes. These findings suggest that LLMs can support clinically meaningful, training-free CKD screening using minimal community-accessible patient features, offering a practical complement to conventional ML methods in real-world screening contexts.
Muhammad Ashad Kabir, S. Munira· Conference on Artificial Int...· 1 citation
This entry-level tutorial aims to equip healthcare professionals with the tools necessary to effectively integrate LLMs into clinical practice, ensuring that these powerful technologies are applied in a safe, reliable, and impactful manner.
Qiao Jin, Nicholas Wan, Robert Leaman et al.· Nature Protocols· 1 citation
This work introduces future querying, a paradigm that probes whether large language models can function as implicit medical world models by evaluating their ability to answer time-indexed clinical queries about a patient's future, and shows that small, locally fine-tuned open-weight models can match or approach larger proprietary systems, making the framework suitable for privacy-preserving, on-premise deployment.
Siri Willems, James Butterworth, L. Goetschalckx et al.· 0 citations
An engineering-oriented, end-to-end roadmap that structures the full lifecycle of clinical language model systems—from model design and domain adaptation to optimization and real-world evaluation is introduced.
PRISM2 demonstrates how language-supervised pretraining provides a scalable, clinically grounded signal for generalizable pathology representations, bridging human diagnostic reasoning and foundation model performance.
E. Vorontsov, George Shaikovski, Adam Casson et al.· Nature Medicine· 3 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.