These findings support the feasibility of leveraging LLMs to enrich EHRs with structured SDOH data, providing a scalable approach for incorporating social context into downstream risk stratification and health outcome prediction.
Abstract
Objectives
Social determinants of health (SDOH) are incompletely captured in structured electronic health records (EHRs) but are frequently documented in unstructured clinical notes. We evaluated large language models (LLMs) for extracting SDOH from clinical text.
Materials and Methods
We constructed an adult sepsis cohort from the Medical Information Mart for Intensive Care-IV (Sequential Organ Failure Assessment ≥2) and analyzed clinical notes from 1 year prior to 30 days following suspected infection. Three instruction‑tuned, decoder‑only LLMs (Mistral‑Instruct‑7B-v0.2, DeepSeek‑R1‑Distill‑Qwen‑14B, and GPT‑oss‑20B) were evaluated using structured prompts with predefined label schemas and few‑shot examples. Performance was benchmarked against a clinically validated annotated dataset and compared with a fine‑tuned encoder‑decoder baseline. Macro‑F1 scores were reported. A gold‑standard Intensive Care Unit (ICU) sepsis subset was independently annotated by 3 reviewers to assess domain‑level performance and ensemble strategies.
Results
Decoder‑only models outperformed the fine‑tuned encoder‑decoder baseline across SDOH domains. GPT‑oss achieved the highest macro‑F1 score (0.79) compared with Flan‑T5‑XXL (0.57). Prompt refinement substantially improved extraction accuracy. Ensemble majority voting increased robustness across domains, while unanimous agreement yielded high precision but limited coverage. In a subsequent mortality analysis, extracted SDOH did not independently predict 30-day mortality, which was instead associated with established clinical and demographic risk factors.
Discussion
Instruction‑tuned decoder‑only LLMs can reliably extract multiclass SDOH from unstructured clinical notes without task‑specific fine‑tuning. Ensemble and agreement‑based strategies provide practical operating points for high‑precision clinical deployment.
Conclusion
These findings support the feasibility of leveraging LLMs to enrich EHRs with structured SDOH data, providing a scalable approach for incorporating social context into downstream risk stratification and health outcome prediction.
The findings support the feasibility of applying LLM-based natural language processing tools in resource-limited, non-English healthcare settings and should assess emerging high-parameter models and explore additional clinical domains.
Breno Gabriel Araújo Sampaio de Jesus, Tomaz Castrillon Figueiredo, Clariele de Almeida Pereira et al.· Cadernos de Saúde Pública· 1 citation
Recent approaches to 30-day hospital readmission prediction rely on pre-trained language models applied to discharge summaries. Although these methods achieve strong performance, they depend on the availability of clinical notes, incur substantial computational costs, and yield representations that lack interpretabilit...
Mohamad Najafi, Hong-Yun Fu, M. Brochhausen et al.· 0 citations
Timely recognition of sepsis remains difficult when early physiological abnormalities are subtle or incomplete. This
study examined whether general-purpose large language models could discriminate sepsis risk from an initial ICU vital-sign
snapshot as effectively as established clinical scoring approaches. We performed...
Anvit More, Vishala Bodetti, Kishan Gor et al.· International Journal for Re...· 0 citations
STRIDE, a machine-learning framework for sepsis detection across seven hospitals with a scalable approach to label quality, outperformed SOFA and Epic on discrimination and showed favorable calibration by Brier score, while retaining strong discrimination among SIRS-positive non-septic encounters.
I. K. Kalyvianakis, C. De Amezaga, E. S. Lee et al.· medRxiv· 0 citations
Findings highlight key trade-offs among symbolic, generative, and instruction-tuned approaches for CGA and 4Ms phenotyping, suggesting that hybrid heuristic-LLM methods offer interpretability and stability, whereas large language models provide greater adaptability for complex clinical narratives.
Sun-Yang Fu, M. J. Kwak, J. Ahn et al.· npj Health Systems· 0 citations
This study provides a systematic comparison of several open-source LLMs on a structured intestinal symptom extraction task and concludes that Qwen3 models offer a favorable balance between accuracy and efficiency, making them suitable for resource-constrained scenarios.
Xin-Yue Zhang, Quan-Yu Wang, Beibei Liu et al.· Journal of Medical Internet...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.