Skip to content
Preprint

CDEG: Learning Decision-Critical Evidence for Long-Horizon Diagnostic Agents

Aug 2026 · 0 citations · 37 references
Computer Science

TL;DR

CDEG, a graph-based framework that learns reusable decision-critical evidence from historical diagnostic trajectories, is introduced, demonstrating that reliable long-horizon diagnosis requires moving beyond trajectory-level experience reuse toward evidence-level learning of the factors that truly shape clinical decisions.

Abstract

Unlike static medical question answering, long-horizon diagnosis captures the sequential nature of clinical practice: evidence is progressively acquired, integrated, and evaluated over multiple rounds of interaction before reaching a final diagnosis. However, existing doctor agents often fail when critical evidence is either not acquired or not adequately incorporated into diagnostic reasoning. Recent agentic approaches attempt to address these failures by reusing historical trajectories or distilled memories. But their diagnostic gains remain constrained because such experience may contain noisy or incidental information and is typically reused without validating which evidence actually drives diagnostic decisions. To address this limitation, we introduce CDEG, a graph-based framework that learns reusable decision-critical evidence from historical diagnostic trajectories. CDEG contrasts successful and failed trajectories from the same case to identify candidate evidence, validates their diagnostic impact through controlled counterfactual interventions, and organizes the resulting diagnosis--evidence--action relations into a structured graph. During inference, CDEG tracks the evolving patient evidence state to retrieve relevant diagnostic relations and selectively guide missing evidence acquisition or overlooked evidence reappraisal. Across in-domain and out-of-distribution benchmarks with multiple doctor agent backbones, CDEG consistently improves diagnostic performance, achieving up to an 11.5% accuracy gain over vanilla agents. These results demonstrate that reliable long-horizon diagnosis requires moving beyond trajectory-level experience reuse toward evidence-level learning of the factors that truly shape clinical decisions.

View source

Similar papers

Preprint Aug 2026

EviDx: Evidence-Aware Active Diagnosis with Scaffolded LLM Agents

EviDx is introduced, an evidence-aware active diagnosis framework that pairs patient-specific diagnostic environments with a clinical diagnostic scaffold and an observer-guided runtime harness that improves diagnostic performance and process stability while revealing model-dependent capability boundaries.

Lihang Zeng, Shao-Ting Zhang, Xiaofan Zhang · 0 citations
#natural language process... Preprint Sep 2026

Can LLMs Value the Right Evidence? Evidence-Value Misalignment in Dynamic Medical Diagnosis

It is demonstrated that verifying evidential support before submission can make diagnostic decisions more reliable, and decouples diagnosis generation from submission through an offline Contrastive Diagnostic Wiki and three online control stages, namely observation management, proposal and witness verification, and dia...

Ke-Hua Feng, Yun-Sheng Lu, Yi-Tong Qiao et al. · 0 citations
Preprint Aug 2026

DiagLoop: A Counterfactual Data Flywheel with Stage-Localized Reinforcement for Diagnostic LLMs

DiagLoop is presented, a counterfactual data flywheel that converts codified physical relations or clinical guidelines, authored once per mechanism family, into training supervision beyond recorded cases, and improves strict path correctness over the strongest conventional baseline.

Jian Zhang, Bing-Yi Wang, Yi-Zhi Liu · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.