Skip to content

2 papers indexed here

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

Preprint Aug 2026

Expert-Guided g-computation with Large Language Models for Estimating Causal Effects on Timings: Applications to Hospital Quality Improvement

Hospital quality improvement (QI) programs routinely face multiple candidate interventions to optimize hospital flow, but existing methods struggle to estimate and rank the causal effects of such interventions. This work focuses on one of the most standard hospital metrics, the average length of stay (LOS), and its causal estimand, the average time saved. To characterize this causal effect, qualitative approaches rely on expert judgment to map patient trajectories, making them susceptible to cognitive biases; quantitative approaches rely on data-driven models, which fail when interventions are hypothetical with no historical data or have complex causal mechanisms that require clinical reasoning rather than data alone. We propose expert-guided g-computation, or egg-computation, which combines the complementary strengths of both approaches by connecting the Gantt charts commonly used to map patient trajectories with the causal DAG literature. We introduce a causal model over Gantt charts and establish identification using a variant of g-computation that seeks expert input only for components unidentifiable from data. To make egg-computation practical, we develop an LLM-assisted pipeline that reliably scales up expert reasoning. In simulations, egg-computation outperforms conventional causal inference methods when patients have diverse causal structures and intervention mechanisms. In a study of eleven candidate QI interventions at an urban safety-net hospital, the LLM pipeline generated graphs and time-saving estimates highly concordant with those of human experts. Beyond healthcare, egg-computation is a broadly applicable framework for estimating the average time saved for candidate interventions whose causal mechanisms can be represented using Gantt charts.

Patrick Vossler, Jialin Ouyang, F. Guo et al. · 0 citations
Open access Aug 2026

Assessing acuity in pediatric emergency department triage: performance of a large language model

Pediatric triage performance varies across emergency departments (ED), contributing to ongoing challenges in pediatric emergency care. There is growing interest in using large language models (LLMs) to support more consistent triage decision-making in children. We evaluated an LLM’s (GPT-5-mini) ability to identify the higher-acuity child from pairs of de-identified clinical notes. Across 228,104 pediatric ED visits, the LLM achieved an overall accuracy of 0.73 (95% CI, 0.73–0.74) in identifying the higher-acuity child, with accuracy improving as acuity differences between visits increased. The LLM was less likely to be correct when the higher-acuity child was older (odds ratio, 0.62, 95% CI, 0.61–0.63) and when the age difference between children was large (0.75, 95% CI, 0.70–0.79). The LLM showed moderate overall accuracy in assessing pediatric acuity and demonstrated a tendency to prioritize younger children, similar to human performance. These findings highlight the need for pediatric-specific LLM evaluation and optimization before clinical use.

Kush Narang, N. Addo, Christopher Y. K. Williams et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.