Author

Saahoon Hong

1 paper indexed here

Fetches their full publication history.

Not the right person? Other researchers publish under this name.

Review Open access Jul 2026

Transformer-Based Language Models for Clinical Decision Support Using Clinical Notes: A Scoping Review

Background/Objectives: This scoping review examined recent evidence on the use of transformer-based language models, encompassing encoder-only architectures (e.g., BERT and its clinical variants) and generative large language models (LLMs; e.g., GPT-4 and Llama), to support clinical decision making from unstructured clinical notes, with implications for behavioral-health services where narrative documentation is central. Methods: Following PRISMA-ScR guidelines, PubMed, PsycINFO, and Web of Science were searched for peer-reviewed studies published between 1 January 2023, and 5 August 2025. Studies applying transformer-based language models to clinical narratives for healthcare tasks and reporting evaluative outcomes were included. We extracted data on clinical tasks, model architectures, enhancement strategies, and evaluation metrics; mapped each study by primary purpose, care setting, and primary model approach; and charted reported validation design, direct human comparison, fairness assessment, workflow evaluation, and clinical deployment. Results: Thirty-six studies were included. Information extraction/de-identification and classification/prediction predominated, whereas summarization/generation was less commonly represented. Model approaches appeared to align with task characteristics: encoder-only and decoder-only systems were frequently used for extraction, encoder–decoder systems for generation, and hybrid or pipeline-based approaches for classification and prediction. Standard task-specific metrics (e.g., F1 and AUROC) predominated, whereas evidence beyond retrospective task performance, including direct human comparison, fairness assessment, workflow evaluation, clinical deployment, and temporal or external validation, was rare. No included study evaluated a transformer-based language model application within a behavioral-health service or behavioral-health workflow. Conclusions: Transformer-based language models have been applied across diverse clinical-note tasks, but the evidence base more strongly supports retrospective task feasibility than transportability, equitable performance, workflow benefit, or safe clinical deployment. Future research should prioritize transparent reference standards, external and prospective validation, clinically meaningful human comparison, and equity-focused evaluation, including direct studies in behavioral-health services.

Saahoon Hong, Hunhui Na · 0 citations