Skip to content
Open access

Validating methods for inferring co-occurring diseases: a flexible framework for simulating synthetic data

Aug 2026 · BMC Medical Research Methodology · Vol 26 · 1 citation · 47 references
Medicine

TL;DR

A four-step framework to generate synthetic data for the simulation-based validation of statistical methods is proposed, broadly applicable beyond the specific medical use case, and relies on careful, domain-informed parameter curation to generate meaningful synthetic datasets.

Abstract

The validation of methods is an integral part of statistical research, defining conditions under which methods yield reliable results. Empirical validation requires a solid data basis to control and manage relevant characteristics like sample size, dimensionality, and underlying dependency structures. Real-world data often fails to meet these requirements, particularly in medical contexts where privacy regulations restrict availability. For this reason, synthetic data is an effective alternative for method validation. However, generating synthetic data is demanding when it must precisely mirror complex dependence structures while simultaneously controlling specific target characteristics. We address the medical context of co-occurring diseases, where symptoms may overlap or conflict. We propose a four-step framework to generate synthetic data for the simulation-based validation of statistical methods. The framework involves: (I) generating patient covariates; (II) connecting this information to predictors for single or joint disease occurrence; (III) transforming predictors into disease probabilities or scores; and (IV) converting these into disease occurrences. Each step offers several alternatives for modeling the overall dependence structure. We apply our framework to a case study of pain-causing diseases which share certain similarities in their clinical presentations, and which can occur either individually or jointly. By employing five combinations of methodological alternatives, we evaluate the approaches’ ability to achieve target characteristics and demonstrate their specific strengths and weaknesses. Matching the data-generating process with the estimation method allows for the successful recovery of input information, such as coefficients and correlations. Target properties like disease prevalence and associations are achieved to varying degrees depending on the methods used. While the proposed theory-driven framework is broadly applicable beyond the specific medical use case, it relies on careful, domain-informed parameter curation to generate meaningful synthetic datasets. Its flexible, adjustable input settings enable researchers to tailor data generation to their precise methodological requirements, providing a controlled basis for simulation-based validation without implying direct clinical inference.

Read PDF

Similar papers

Review Sep 2026

A Conditional-Distribution Framework for Validating Synthetic Multivariate Data

Statistical validation of synthetic multivariate data requires assessing whether a generator preserves the joint dependence structure of the target population without merely reproducing observed records. We develop a model-agnostic framework based on full conditional distributions. For each coordinate, we normalize the...

H. Dahal, Ishanu Chattopadhyay · 0 citations
Open access Aug 2025

Can synthetic data reproduce real-world findings in epidemiology? A replication study using adversarial random forests

Abstract Background Synthetic data hold substantial potential to address practical challenges in epidemiology due to restricted data access and privacy concerns. However, many current methods suffer from limited quality, high computational demands, and complexity for non-experts. Furthermore, common evaluation strategi...

J. Kapar, K. Günther, L. Vallis et al. · 1 citation
Preprint Aug 2026

A Statistical Framework for Data-Driven Discovery of Differential Performance in Clinical Risk Prediction Models

The unfairness tree (utree) is proposed, a data-driven recursive partitioning framework for identifying subgroups with differential model performance that exhibits nominal empirical type I error rates and good ability to detect, quantify, and characterize performance discrepancies defined by higher-order variable inter...

A. Neher, Julian Wolfson · 0 citations
Open access Sep 2026

Leveraging Synthetic Clinical Data for Validation and Operational Readiness in Clinical Trials

Background/Objectives: Obtaining timely access to detailed clinical trial data is not always straightforward. Privacy requirements, governance processes, and study-specific eCRF configurations can delay access, particularly during study start-up, when teams need realistic data to develop and test validation rules, repo...

Szymon Musik, Jacek Zalewski, Julia Jurkowska et al. · 0 citations
Open access Aug 2026

Maria-X: A Multimodal Transformer Aware of Missing Data for Reliable Diagnosis and Prognosis of Alzheimer’s Disease

Missing values are a persistent issue across real-world datasets, and they often undermine the reliability of models built on top of that data. Most existing solutions have been developed and tested under the MCAR assumption a condition that is uncommon in actual data collection settings and therefore leaves open quest...

Lakshmiprasannakumar Vemavarapu, Chandra Sekhar Sanaboina · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.