Skip to content
Open access

SynTrustBench: An Evidence-Gated and Executable Benchmark for Trustworthiness Claims in Synthetic Clinical Data

Aug 2026 · medRxiv · 0 citations
Medicine

TL;DR

SynTrustBench is introduced, an evidence-gated and executable benchmark for evaluating trustworthiness claims across five non-compensable dimensions: fidelity, clinical utility/validity, privacy, equity, and robustness/generalization.

Abstract

Synthetic clinical data are increasingly used for healthcare machine-learning development, model validation, data sharing, and predeployment testing, yet such data often claim to be trustworthy after passing a limited collection of realism tests. A synthetic dataset may indeed claim statistical similarity while leaking training membership, erasing rare subgroups, failing on held-out real patients, or lacking sufficient artifacts for reproduction. We introduce SynTrustBench, an evidence-gated and executable benchmark for evaluating trustworthiness claims across five non-compensable dimensions: fidelity, clinical utility/validity, privacy, equity, and robustness/generalization. Its Evidence Assessment component audits published reports and produces a five-element Evidence Maturity Profile (EMP) together with a separate evaluability gate. Its executable structured-tabular protocol accepts frozen real training data, held-out real test data, a synthetic table, and a declarative configuration; computes dimension-specific metrics and uncertainty; and produces subgroup results, failure flags, benchmark cards, and provenance manifests. In a frozen pilot audit of 30 reports, 17 of 30 quantitatively evaluated privacy, 2 of 30 documented a formal privacy guarantee to the audit threshold, 2 of 30 evaluated equity, 12 of 30 evaluated robustness, and only 4 of 30 passed the evaluability gate. The executable implementation operationalizes the same dimensions through distribution and dependency checks, frozen train-on-real/test-on-real (TRTR) and train-on-synthetic/test-on-real (TSTR) utility, empirical privacy attacks, subgroup analysis, perturbation testing, and a controlled failure-injection harness. SynTrustBench does not certify clinical safety or collapse trustworthiness into a single score. Instead, it provides an inspectable predeployment contract for identifying what was evaluated, what failed, what remains unknown, and whether evidence is sufficiently complete and reproducible for comparison or downstream healthcare AI use.

Read PDF

Similar papers

#machine learning Review Aug 2026

CoMedBench: A Multi-Source Benchmark of Synthetic Medical Data Fidelity and Downstream Utility

CoMedBench is introduced, a reproducible benchmark that evaluates a family of generators under a common clinical-validity framework and one shared training and evaluation engine, spanning static tabular and temporal downstream tasks on established critical-care datasets.

Akanta Das, Farhad Al-Amin Dipto, Mrinmoy Sarkar Anto et al. · 0 citations
Book Open access Aug 2026

LiveMedBench: A Contamination-Limited Medical Benchmark for LLMs with Automated Rubric Evaluation

The deployment of Large Language Models (LLMs) in high-stakes clinical settings demands rigorous and reliable evaluation. However, existing medical benchmarks remain static, suffering from two critical limitations: (1) data contamination, where test sets inadvertently leak into training corpora, leading to inflated performance estimates; and (2) temporal misalignment, failing to capture the rapid evolution of medical knowledge. Furthermore, current evaluation metrics for open-ended clinical reasoning often rely on either shallow lexical overlap (e.g., ROUGE) or subjective LLM-as-a-Judge scoring, both inadequate for verifying clinical correctness. % To bridge these gaps, we introduce LiveMedBench, a continuously updated, contamination-limited, and rubric-based benchmark that weekly harvests real-world clinical cases from online medical communities, ensuring strict temporal separation from model training data. We propose a Multi-Agent Clinical Curation Framework that filters raw data noise and validates clinical integrity against evidence-based medical principles. For evaluation, we develop an Automated Rubric-based Evaluation Framework that decomposes physician responses into granular, case-specific criteria, achieving substantially stronger alignment with expert physicians than LLM-as-a-Judge. % To date, LiveMedBench comprises 2,756 real-world cases spanning 38 medical specialties and two languages, paired with 16,702 unique evaluation criteria. Extensive evaluation of 38 LLMs reveals that even the best-performing model achieves only 39.2%, and 84% of models exhibit performance degradation on post-cutoff cases, confirming pervasive data contamination risks. Error analysis further identifies contextual application---not factual knowledge---as the dominant bottleneck, with 35-48% of failures stemming from the inability to tailor medical knowledge to patient-specific constraints. The code and data are available at https://github.com/ZhilingYan/LiveMedBench/ LiveMedBench.

Zhiling Yan, D. Song, Zhe Fang et al. · 0 citations
Open access Jul 2026

Addressing benchmarking gaps in large language models for health and medicine with dynamic red-teaming

A Dynamic, Automatic and Systematic red-teaming audit framework that continuously stress-tests LLMs for health across four safety-critical axes: robustness, privacy, bias and hallucination, which provides a scalable framework for surfacing latent risks before such systems are deployed in consumer-facing health assistants and broader clinical workflows.

Jiazhen Pan, Bailiang Jian, Paul Hager et al. · 0 citations
Review Aug 2026

CliniCARE-Bench: Clinical Calibrated Audit of Medical Reasoning in EHR

CliniCARE-Bench is the first deployment-oriented clinical-agent benchmark to jointly evaluate real longitudinal EHR investigation, claim-level evidence grounding, governing-policy use, process adherence, and calibrated abstention within a common patient-level adjudication framework.

Veronica Chatrath, Bryan Zhu, George Pu et al. · 1 citation
Jul 2026

ClinLens: Towards Long-Horizon Coding Agents for Longitudinal Multimodal Clinical Data Science

CLINLENS is introduced, a benchmark of 200 executable tasks over five linked MIMIC resources spanning structured electronic health records, notes, electrocardiograms, chest radiographs, and echocardiograms, which exposes a substantial gap between runnable submissions and correct clinical analyses.

Yuan Zhu, Ethan B. Liu, Frank Nie et al. · 0 citations
#artificial intelligence Preprint Aug 2026

CareGraph: An Auditable Hybrid AI Framework for Evidence-Grounded Personalized Longitudinal Health Intelligence

CareGraph is an auditable hybrid AI framework that converts heterogeneous records into prioritized trends, missing context indicators, bounded next steps, discussion questions, and provenance linked explanations that offers a safety bounded foundation for intelligent personalized health systems.

Prof. S. B. Ghawate, Tanvi R. Patil · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.