Skip to content
Open access

From Clinical Free Text to Auditable Concepts: An Agentic Framework for Interpretable Prediction

Sep 2026 · medRxiv · 0 citations
Medicine

TL;DR

An agentic workflow that takes a prediction task and a raw text corpus as input and produces an auditable feature layer that achieves predictive performance comparable to direct BioClinicalBERT prediction on both tasks while additionally providing explicit, interpretable, and traceable task-specific concepts.

Abstract

Across application domains, predictive signals often sit in unstructured free text rather than structured fields, yet turning that text into useful and interpretable features is difficult. Running large language models (LLMs) over an entire corpus is costly and hard to reproduce, while end-to-end text representations can rely on surface cues that are difficult to inspect. We present an agentic workflow that takes a prediction task and a raw text corpus as input and produces an auditable feature layer. The first two agents use an LLM to derive a task-specific predictor taxonomy and weakly label a bounded text sample; routed local extractors then process the corpus, and a deterministic builder aggregates the evidence into a dynamic, longitudinal concept bottleneck. We evaluate the framework on medication discontinuation in a longitudinal oncology cohort and 30-day readmission in MIMIC-IV. With gradient boosting, the longitudinal bottleneck increases area under the receiver operating characteristic curve (AUROC) over coarse concept buckets from 0.700 to 0.761 for medication discontinuation and from 0.576 to 0.609 for readmission. The proposed framework achieves predictive performance comparable to direct BioClinicalBERT prediction on both tasks while additionally providing explicit, interpretable, and traceable task-specific concepts. LLM use is confined to a bounded weak-labeling stage costing $24.00 and $23.39, respectively, compared with projected costs of $10,648 and $11,519 for exhaustive sentence-level LLM processing of the full corpora, demonstrating the substantial cost efficiency of the proposed agentic system.

Read PDF

Similar papers

Preprint Aug 2026

Making Clinical Language Models Auditable: Concept-Guided Fine-Tuning for Robust Prediction

CAST (Concept-guided Artifact Suppression Tuning), an SAE-based framework for auditable clinical text classification, improves over its corresponding fine-tuned encoder baselines and remains competitive with strong LLM baselines, while producing a feature-level audit trail of the clinical concepts that support each pre...

Jin Mu, Guan-Hua Chen · 0 citations
Book Open access Aug 2026

OneEHR: Reproducible and AI Agent-Ready Longitudinal EHR Analysis Toolkit

Electronic health records support a wide spectrum of clinical prediction and decision-support studies, but reproducible EHR research now requires more than training a single predictive model. As the field expands from machine learning and deep learning to LLM-based and agentic AI, differences in cohort construction, te...

Yinghao Zhu, Zi-Xiang Wang, Lei Gu et al. · 0 citations
#artificial intelligence Preprint Sep 2026

Unknown is not normal: separating language-model extraction from rule-based decision logic for clinical risk scores

Large language models (LLMs) are increasingly used to compute clinical risk scores from free-text notes. Notes are often incomplete, and treating undocumented findings as normal can silently misclassify patients. We test whether separating three-state extraction (present, absent or unknown, by an LLM) from decision log...

Nicolás Vera Zúñiga · 0 citations
#small language model Preprint Aug 2026

Future Querying: Can LLMs Serve as Implicit Medical World Models?

This work introduces future querying, a paradigm that probes whether large language models can function as implicit medical world models by evaluating their ability to answer time-indexed clinical queries about a patient's future, and shows that small, locally fine-tuned open-weight models can match or approach larger...

Siri Willems, James Butterworth, L. Goetschalckx et al. · 0 citations
Preprint Aug 2026

Explainable Transformer Models for Clinical Prediction Tasks on Structured Electronic Health Records

BERT-LER is presented, a BERT-style model for coded EHR timelines pretrained and fine-tuned from a de-identified EHR dataset of 75 million patients, that encodes laboratory test results as discrete tokens while retaining graded information through percentile-based binning, paired with Integrated Gradients for token-lev...

Jun-Ni Du, Lukas Adamek, Maxim A Kryukov et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.