Clinical question answering over electronic health records (EHRs) increasingly relies on large language model (LLM) agents that retrieve structured patient data through external tools. Published benchmarks, however, evaluate these systems at a single patient-population size, and rarely measure the effect of backend representation from that of the retrieval interface design. This paper compares six retrieval configurations that vary along two axes: backend (a property graph database, a relational database and a dense vector index) and interface design (curated domain-specific tool calls, model-generated queries, full-text search, and single-shot dense retrieval). The evaluation covers a 334-question bank spanning six categories (simple lookup, multi-hop, temporal, cohort, reasoning, and unanswerable), instantiated at three nested population scales: 200, 2000, and 20,000 alive patients from a single Synthea cohort. Four models are compared: Claude Haiku 4.5, Qwen 2.5 72B, Llama 3.1 8B, and Llama 3.3 70B, spanning closed-frontier and open-source alternatives. Curated tool-calling configurations improve accuracy over retrieval-augmented baselines for capable models, but reduce accuracy for a small open-source model due to function-calling protocol failures. We report how accuracy, latency, and cost evolve with each approach, model size, and cohort size, supported by paired statistical tests and confidence intervals. All benchmark components, databases, and evaluation code are publicly available.
Leonidas Anagnou, Andreas Vezakis, Ioannis A Vezakis et al.· Future Internet· 0 citations
Accurate segmentation of abdominal organs in Computed Tomography (CT) underpins radiotherapy planning, surgical planning, and disease monitoring. Existing benchmarks rank architectures by a single aggregate Dice score, without per-organ statistical testing or boundary-sensitive metrics, even though models are chosen organ by organ for clinical use. We benchmark ten architectures spanning convolutional, attention-based, transformer, and state–space (Mamba) families on the AMOS CT dataset under one identical nnU-Net-style pipeline; we report per-organ Dice, 95-percentile Hausdorff Distance (HD95), and Normalised Surface Dice, with pairwise significance tested on an independent external dataset (TotalSegmentator). A competitive cluster of convolutional and Mamba models leads; rankings are stable on large organs but reshuffle by 10–13% on the small, geometrically complex ones, and boundary fidelity separates the models into tiers that the Dice ranking hides. This ordering largely holds on the external set (Spearman ρ=0.84). Selecting a model on aggregate Dice alone is therefore unsafe for organ-specific clinical tasks: per-organ overlap and boundary metrics should be the primary acceptance criteria for selecting a model before clinical deployment.
Alexandros Barmperis, Olga Menegaki, Anna Panagiotakopoulou et al.· Journal of Imaging· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.