Skip to content
Conference

Clinical Reasoning-Augmented Prediction: Agentic Verification Improves Reliability of VTE Detection from Radiology Reports

Jul 2026 · International Conference on Digital Health · pp. 1-11 · 0 citations · 24 references

Abstract

Automated extraction of clinical outcomes from radiology reports is widely used for cohort construction; extracted outcomes also serve as ground truth labels to evaluate clinical Machine Learning (ML) models. However, conventional Natural Language Processing (NLP) classifiers rely primarily on statistical pattern recognition and often fail when clinical meaning depends on context, such as negation, chronic versus acute disease characterization, and mention of past or family history. These contextual failures reduce reliability even when the reported accuracy appears high. We propose a clinical reasoningaugmented architecture for venous thromboembolism (VTE), including the detection of acute deep vein thrombosis (DVT) and acute pulmonary embolism (PE) from radiology reports, comprising multiple stages of AI model prediction and clinical interpretation. A fine-tuned discriminative Large Language Model (LLM) classifier (Qwen-7B with sequence classification head, LoRA/QLoRA) produces an initial VTE prediction with associated class posterior probabilities from the full report text. A second stage, implemented as an agentic verification LLM trained via reinforcement learning (GRPO), re-reads the report and verifies whether the predicted label is clinically justified. The verifier focuses on diagnostic sections (IMPRESSION, FINDINGS, CONCLUSION) and applies contextual constraints including negation detection, temporal qualifiers (acute vs. chronic), historical mention identification, internal report consistency checks, and evidence-grounded clinical rules. We evaluate VTE cases from multiple institutions across 144 VA hospitals (VA), University of Maryland Medical Center (UMMC), and a public Stanford INSPECT dataset. The two-stage system consistently outperforms baseline LLM classifiers, achieving accuracies of 0.9957 (VA DVT), 0.9950 (UMMC Reader1), 0.9942 (UMMC Reader2, including challenging acute vs. chronic cases), and 0.9905 (Stanford PE), while substantially reducing misclassification rates, particularly for clinically ambiguous reports.

View source

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.