AEGIS advances generative AI that is fit for decision-making by grounding models in interventions, counterfactuals, and calibrated uncertainty. The workshop, held during ACM KDD 2026 conference, convenes researchers and practitioners in causality, LLMs, and prescriptive analytics to address when and how generative systems should recommend actions in healthcare and public policy. We invite methods that couple structural causal reasoning with LLMs, diffusion and sequence models; techniques for off-policy evaluation, dynamic treatment regimes, and feedback-aware learning; semi-synthetic benchmarks and governance practices; and measures beyond predictive fidelity, including policy regret, counterfactual calibration, and safety constraints. The program will feature a keynote talk and peer-reviewed presentations. By aligning with KDD's emphasis on trustworthy, scalable AI, AEGIS endeavors to establish shared evaluation protocols and artifacts that make prescriptive models reliable under distribution shift, so recommendations remain robust and accountable from development to deployment across settings and populations.
M. Prosperi, Yi Guo· Proceedings of the 32nd ACM...· 0 citations
Abstract Objectives To evaluate the real-world performance of a transformer-based natural language processing (NLP) system for extracting Social Determinants of Health (SDoH) from clinical notes, using survey-based SDoH measures a reference comparators. Materials and Methods This study was conducted at the University of Florida Health in adults with at least 2 clinical encounters in the prior year. A research survey was completed by 1001 participants; sampling targeted 50% Black patients to support subgroup analyses. Comparative analyses were restricted to the 414 participants who also had Epic SDoH survey data and clinical notes available for NLP extraction. We compared concepts extracted by the SOcial DeterminAnts (SODA) NLP pipeline against the independently administered research survey, which served as the primary reference standard, and against the structured Epic-embedded SDoH questionnaire. Nine domains were evaluated: abuse, alcohol use, drug use, education, financial constraints, housing, physical activity, social cohesion, and transportation. Sensitivity, specificity, positive predictive value, negative predictive value, and F1 scores were calculated by domain. Results The NLP pipeline more consistently aligned with negative survey responses than with patient-reported social needs, although performance was lower in some domains, especially alcohol use and financial constraints. Sensitivity was higher only for alcohol use (55%); the lowest values were for abuse (5%), drug use (0%), and financial constraints (16%). These results cannot be attributed to SODA extraction alone. They reflect some combination of social information not being recorded in clinical notes, content that was recorded but not extracted, and mismatches between extracted concepts and survey definitions, and the present analysis cannot separate these contributions. The 2 surveys agreed only modestly with each other, so no single instrument provides a definitive ground truth. Conclusion The pipeline more consistently aligned with negative survey responses than with patient-reported social needs. Because the 2 surveys agreed only modestly, the reference itself is imperfect, and apparent NLP performance depends in part on which survey is used as the comparator. Apparent gaps in NLP performance reflect both how social risks are recorded in clinical notes and how patients disclose them across different survey settings, in addition to limits of the extraction pipeline. Improving documentation practices, integrating locally tuned large language models, and monitoring subgroup performance may all be needed to make SDoH identification tools reliably detect social needs across patient populations.
Xiangren Wang, Jingchuan Guo, Yi Guo et al.· JAMIA Open· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.