Case2Flow is introduced, a task designed to retrieve the most relevant guideline flowchart for a given patient case from a collection of guideline documents and CRISP, a training-free scoring method that sharpens late-interaction retrieval by suppressing uninformative patches, discounting ambiguous token matches, and incorporating bidirectional query-image alignment.
Abstract
Medical guidelines encode rich, evidence-based decision logic, yet the specific decision artifact a clinician needs is hard to locate within a guideline, let alone across guidelines covering plausible diseases and treatments. While guideline passages have supported end-to-end question answering, flowcharts remain largely underused in decision support despite their ability to encode actionable clinical pathways. We therefore introduce Case2Flow, a task designed to retrieve the most relevant guideline flowchart for a given patient case from a collection of guideline documents. To support it, we construct FlowAtlas, a curated corpus of 202 flowcharts extracted from 2,080 medical guidelines, together with a pipeline that synthesises 1,911 aligned case-flowchart pairs. Our evaluation of multimodal retrieval methods reveals systematic failure modes, including overreliance on keywords and spurious token-patch matches induced by uninformative background regions in flowcharts. Motivated by this, we propose CRISP, a training-free scoring method that sharpens late-interaction retrieval by suppressing uninformative patches, discounting ambiguous token matches, and incorporating bidirectional query-image alignment. CRISP improves Recall@1 by up to 18.71 percentage points, while a blinded physician assessment on published case narratives provides preliminary feasibility evidence beyond synthetic queries.
Understanding medical search intent is critical for improving biomedical information retrieval, particularly for identifying evidence-seeking queries. However, existing resources are limited in scale and lack structured annotation aligned with clinical evidence needs. We present MedQueryIntent, a large-scale benchmark dataset consisting of a human-annotated real-world query set and an LLM-augmented evidence query set. The human-annotated portion, MedQueryIntent-Real, contains 12,000 medical queries collected from Trip and PubMed and is labeled with a hierarchical taxonomy distinguishing evidence-seeking and non-evidence-seeking intent. The annotation follows structured guidelines based on clinical research frameworks (e.g., PICO, PEO, PCC), ensuring consistent and interpretable labeling. To address the limitations of short and ambiguous real-world queries, we further construct MedQueryIntent-Synth, an LLM-augmented evidence query dataset containing 7,633 unique queries, generated by transforming systematic review content into natural query forms via controlled prompting. Together, the two resources form the complete MedQueryIntent benchmark with 19,633 queries. We benchmark multiple models, including biomedical encoders and large language models, demonstrating that MedQueryIntent provides a challenging and realistic testbed. Results show that incorporating MedQueryIntent-Synth improves classification performance and robustness. Our dataset supports research in medical query understanding and downstream applications such as evidence retrieval and clinical decision support. The dataset is publicly available at https://github.com/yingchengsun/MedQueryIntent.
S. Schnell, Supriya Kottam, Yingcheng Sun· International Conference on...· 0 citations
Measured extraction performance varied substantially by matching definition, whereas exact-link decisions varied with the availability of matched terminology evidence, support separate evaluation of extraction, retrieval, evidence-grounded linking, and extension candidacy.
Objective: To characterize the kinds of internal documentation inconsistencies a general-domain large language model (LLM) can surface from real-world discharge summaries, and to identify recurring failure modes that limit reliability at scale. Materials and Methods: We applied a two-stage LLM pipeline---open-ended candidate identification (Gemini 2.5 Pro) followed by context-grounded verification (Gemini 2.5 Flash)---to 3,000 randomly sampled MIMIC-IV-Note discharge summaries. A subset of the pipeline output was then reviewed manually by clinical experts. Results: Our pipeline surfaced 3,460 candidate inconsistencies, affecting 69.7% of admissions. Representative examples spanned demographics, allergies, procedures, diagnoses, laboratory, medications, and care-planning domains, with direct implications for clinical reasoning or patient safety. Expert review also revealed recurring failure modes that arise when verification requires temporal reasoning, evolving-diagnosis context, or knowledge of outpatient-prescribing conventions the model does not natively possess. Discussion: Detection is highly context-dependent: many flagged pairs require anchoring each statement to its source section and clinical domain, then assessing whether the conflict reflects a true contradiction or missing context. We propose a graded ontology spanning strict contradiction and ambiguity, with a schema characterizing each flagged case by category, section, domain, and inconsistency axis. Conclusion: This formative study establishes a methodological foundation and conceptual framework to guide subsequent validated, large-scale EHR-inconsistency analysis.
A benchmark score credits final answers, but not the route by which an item can be answered. In medical multimodal multiple-choice questions (MCQs), this distinction matters because a correct answer can be supported by the intended image finding or by benchmark-preserved cues in the wording of answers, non-visual clinical text, visible image text, artificial annotations, or device/context artifacts. We call the resulting score-level overinterpretation reasoning inflation. Here, a route is an observable input path that can support answer selection, not a claim about the model's hidden cognition. Across six medical multimodal MCQ datasets, we separate candidate cues from behavioral evidence through prompt- and image-side audits, modality ablations, and matched repairs that preserve the medical target and answer key. In a 13-configuration open-model panel, full-input accuracy is 62.63%, while text-only and options-only settings achieve 53.96% and 29.71%, respectively. Removing length-gap, absolute/conspicuous, and spatial/prepositional cues lowers accuracy by 6.58, 3.50, and 4.77 percentage points. We also construct MedQA-MM, a 1,000-item shortcut-mitigated subset, where text-only and options-only accuracy fall to 5.21% and 12.33%. This does not imply that models never use images; it shows that medical image-reasoning claims require route-level evidence.
Ben-Lu Wang, Yi-Fan Zhang, Jia-Qing Yu et al.· 0 citations
The observed performance gains with RAG support the use of LLMs augmented with quality-assured external knowledge in the German healthcare context, while strong performance of open-weight models suggests potential for on-premises deployment in privacy-sensitive clinical environments.
Johannes Schwietering, G. Lichtner· BMJ Health & Care Informatic...· 0 citations
Clinical Practice Guidelines (CPGs) encode evidence-based clinical knowledge but are primarily distributed as unstructured PDF documents, making them inaccessible to automated clinical decision support (CDS) systems. This paper proposes a service-oriented architecture that formalizes the IMSS Clinical Practice Guideline for Major Burn Management (IMSS-375) as a versioned REST/JSON microservice. Seven clinical decision endpoints are defined, each encapsulating a specific GPC recommendation: burn classification, initial assessment, fluid resuscitation, pain management, infection prevention, nutritional support, and transfer criteria, following HL7 FHIR R4 interoperability standards. A mapping between GPC clinical rules (including the Parkland formula, Benaim scale, Curreri formula, and Baux prognostic index) and structured JSON request/response schemas is presented and evaluated against related formalization approaches and verified through structured schema-level invocations against a representative clinical scenario, including a detailed comparison of Mexico IMSS-375 standard properties against HL7 FHIR and OpenEHR. A deployment architecture is described covering hospital-level integration, a centralized service layer, and a non-relational persistence tier based on document-oriented storage for unstructured clinical data.
José de Jesús Álvarez Ramírez, Rocío Maciel, V. Larios· IEEE Latin America Transacti...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.