Skip to content
Review

Clinical evaluation of a vision-language model for optimizing triage and clinical workflows in critical care.

Aug 2026 · International Journal of Medical Informatics · Vol 221, pp. 106673 · 0 citations · 37 references
Medicine

TL;DR

The VLM demonstrated reliable clinical interpretation and an acceptable safety profile, however its integration into clinical workflows for early recognition of physiological deterioration and patient acuity assessment requires further rigorous evaluation and comparison to currently used track-and-trigger systems and patient monitoring methods.

Abstract

Objective

Clinical deterioration in hospitalized patients is often preceded by subtle, dynamic physiological changes that are difficult to detect using intermittently charted electronic health record (EHR) data. Our objective was to evaluate the reliability, interpretability, and clinical relevance of a Vision‑Language Model (VLM)-based triage framework that analyzes physiological trend images, by comparing VLM-generated outputs with attending physician assessments as the expert clinical comparator.

Materials And Methods

We conducted a single-center expert agreement pilot study including 100 adult patients with two hours of dynamic monitoring data across four vital signs (SpO2, RR, HR, BP). A structured prompt was developed using the Gemini 2.5 Flash model. Two independent reviewers assessed VLM outputs for clinical interpretation, artifact detection, and triage classification. We evaluated reviewer agreement using percent agreement and weighted Cohen's κ. A secondary risk-oriented analysis measured classification concordance and over-triaged cases as lower-risk classifications, while cases underestimating patient acuity were designated as higher-risk misclassifications.

Results

The VLM demonstrated moderate to substantial agreement in triage classification with the attending physician and a low rate of under-triage. The VLM assigned the same patient acuity category as attending physician in 75 % of cases and underestimated acuity in 7 % of cases, compared with 14 % underestimation by the physician in training.

Discussion

VLMs extend generative artificial intelligence capabilities by enabling image‑grounded clinical reasoning and offer signal‑processing capabilities for interpreting time‑stamped physiological trends.

Conclusion

The VLM demonstrated reliable clinical interpretation and an acceptable safety profile, however its integration into clinical workflows for early recognition of physiological deterioration and patient acuity assessment requires further rigorous evaluation and comparison to currently used track-and-trigger systems and patient monitoring methods.

View source

Similar papers

Open access Aug 2026

AI-assisted rheumatology triage changes with referral framing.

OBJECTIVES Two in three US physicians now use healthcare AI, and large language models (LLMs) are entering the triage workflows that determine which patients reach rheumatology and how quickly. We aimed to test whether nine prespecified cues in referral notes, patient descriptions and demographics shift AI-assisted triage decisions when the underlying clinical information is unchanged. METHODS We conducted a controlled, physician-validated experiment across 30 physician-authored rheumatology vignettes, each independently rephrased three times (90 case variants). We tested five LLMs from three providers - Anthropic, Google, and OpenAI - under nine dimensions spanning demographics, clinical context, and communication framing, with 57 controlled contextual modifications, four system-prompt personas, and five repetitions per cell, yielding more than 200,000 model queries. Sixteen clinical fields were graded against physician-validated ground truth, with excellent inter-rater agreement (Fleiss' kappa=0.92). RESULTS Baseline composite concordance with expert ground truth was high at 0.869. We had expected sociodemographic cues to be the strongest source of distortion. Instead, the largest shifts came from how the case was framed and described. When patients were described as anxious, models attributed symptoms to psychological rather than organic causes nearly three times as often as at baseline (13.1% vs 4.5%; odds ratio 3.2 versus stoic framing), a shift that risks relabeling organic disease as functional. Clinician anchoring in the referral note reduced concordance, consistently across models and rephrasings and significantly for acuity (dismissive anchor, vignette-level p = 0.01), mainly by downgrading urgency. In contrast, race or ethnicity, socioeconomic status, and language barrier produced no detectable effect, including in mixed-effects models that accounted for repeated vignette use. CONCLUSION Although baseline concordance with specialist ground truth was high, it was readily disrupted by how referral notes were worded and how patients described their symptoms, not by patient demographics. Before AI-assisted triage enters rheumatology referral pathways, systems should separate objective clinical evidence from interpretive framing, and urgency and psychological attribution should be treated as auditable safety signals.

Mahmud Omar, Mohammad E. Naffaa, R. Agbareia et al. · 0 citations
Review Open access Aug 2026

An Electronic Health Record–Integrated, Large Language Model–Powered Tool to Triage Surgical Patients

Key Points Question Can surgical patient triage be automated using a large language model (LLM) agentic workflow? Findings In this quality improvement study, the LLM tool recommended hospitalist consultation for nearly a quarter of the 6193 triaged cases. The tool achieved 94% sensitivity and 74% specificity, and post hoc medical record review suggested that most discrepancies reflected modifiable gaps in clinical criteria, institutional workflow, or physician practice variability, rather than LLM misclassification. Meaning The findings of this study suggest that an LLM-powered human-in-the-loop agentic workflow could accurately triage surgical patients for a surgical comanagement service.

Janelle B. Wang, T. Keyes, April S. Liang et al. · 0 citations
Review Aug 2026

RESPClinBench: Benchmarking Multimodal Clinical Decision-Making and Longitudinal Disease Management in Respiratory Specialty Care

Background: Respiratory specialty care requires multimodal interpretation, longitudinal risk assessment, guideline-concordant intervention, and whole-course management, which are poorly represented by examination-oriented medical benchmarks. Objective: To develop RESPClinBench, a real-world scenario-based benchmark for respiratory clinical decision-making, and evaluate seven contemporary large language models across AECOPD-PIM and PNBIM. Methods: RESPClinBench cases were adapted from de-identified respiratory clinical data. Three attending-level respiratory physicians revised cases, reference answers, and atomic clinical-action points, while one senior respiratory specialist performed cross-review and final adjudication. AECOPD-PIM comprised 427 open-ended COPD cases, and PNBIM comprised 196 multimodal pulmonary nodule cases combining chest CT with structured clinical information. Seven models generated 4,361 responses through standardized API inference with temperature 0 and a maximum output length of 8192 tokens. An automated framework calculated the final score as the arithmetic mean of atomic-action recall and rubric-based LLM-as-a-Judge assessment. Results: Across 623 cases, the mean final score was 68.58. Qwen3.6-27B ranked first overall at 71.22, Qwen3.5-397B-A17B led PNBIM at 72.48, and Qwen3.6-27B led AECOPD-PIM at 71.11. Imaging hallucination and serious medical risk occurred in 31.85% and 8.16% of PNBIM responses; medication-safety risk and serious medical risk occurred in 26.93% and 1.44% of AECOPD-PIM responses. Conclusions: RESPClinBench identifies task-specific limitations in multimodal pulmonary nodule assessment and longitudinal COPD management. Combining explicit clinical-action coverage, holistic evaluation, and independent safety flags provides a clinically grounded basis for model selection and prospective validation.

Mou-Xiao Bian, Zhi Chen, Ruiyao Chen et al. · 0 citations
Aug 2026

Comparative evaluation of large language models and clinicians in real-world glaucoma clinical reasoning.

Large language model-based AI systems produced structured glaucoma-related reasoning with performance that overlapped with attending ophthalmologists but did not establish clinical equivalence, but may have potential as supervised decision-support and educational tools.

Hou-Fa Yin, Lixia Shen, Haiyan Cai et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.