Skip to content

[Comparing synthetic data generation methods in pharmacoepidemiology: reconciling reproducibility with privacy protection].

2026 · Epidemiologia & Prevenzione · Vol 50 3, pp. 279-289 · 0 citations
Medicine

TL;DR

This study compares two methods of generating tabular synthetic data: synthpop, based on transparent and interpretable inferential methodologies, and Conditional Tabular-Generative Adversarial Networks (CT-GANs), which leverage deep learning approaches to reproduce complex multivariate distributions.

View source

Similar papers

Open access Jul 2026

CLEO closed loop framework for synthesizing medical privacy preserving tabular data

Findings indicate that CLEO provides a controllable and empirically auditable framework for supporting cross-institutional research when real-world medical data cannot be directly aggregated.

Siqi Wang, Jianfeng Wang, Xiaochun Cheng et al. · 0 citations
Open access Aug 2026

Synthetic Longitudinal Tabular Data Generation via Copula

Synthetic data generation is increasingly used to enable data sharing and secondary analysis while protecting participant privacy, particularly for longitudinal tabular health data, where repeated measures per subject create within-subject dependence that most synthetic data methods are not designed to preserve. Existing generative methods, particularly generative adversarial network (GAN)-based approaches, can model complex distributions, but their estimated dependence structures are often difficult to interpret and their performance may be unstable or prone to overfitting in modestly sized datasets. Here we show that eCDF-copula, a statistically rooted approach using the empirical cumulative distribution function (eCDF) and copula modeling, preserves within- and between-visit dependence structure. To handle pervasive missing data, we propose a two-stage strategy combining multiple imputation with copula-based synthesis, enabling a variance decomposition that quantifies replication variability across methods. We benchmarked the proposed approach against four established methods on two longitudinal clinical datasets spanning markedly different sample sizes (n = 120 vs. n = 3, 612). eCDF-copula achieved resemblance and utility exceeding those of state-of-the-art synthetic data methods, while maintaining comparable privacy.

Hanchang Cai, Wen-Shan Yu, Ruijin Lu et al. · 0 citations
Review Open access 2026

Synthetic Data Quality Evaluation in Generative AI: Current Trends, Challenges, and Future Directions for Social Science Research

The findings indicate that while modern generative models can produce highly realistic and analytically useful datasets, persistent challenges remain, including the lack of standardized benchmarking protocols, utility–privacy trade-offs, privacy leakage risks, bias amplification, limited explainability, and governance concerns.

N. Emran, Ruhaila Maskat, Abdulrazzak Ali · 0 citations
Preprint Aug 2026

Novel Knowledge-Guided Generative Methods for Synthetic Transcriptomic Data

As biomedical research increasingly relies on data-intensive tools, the quality and utility of datasets are critical. Challenges such as imbalances, biases, and ethical or legal constraints often limit access to high-quality data. Synthetic data generation can help overcome these limitations. Here, we present a comparative analysis of generative models for transcriptomic data, investigating strategies to incorporate prior biological knowledge via gene graphs. This ensures that synthetic data capture real-world gene patterns, maintaining their usefulness for downstream tasks. In particular, we introduce and benchmark three variants of the Generative Adversarial Network. Among the alternatives, MK-TGAN - an innovative multi-kernel, Graph Neural Network-based model - stands out for its performance in terms of both the realism and utility of the generated data. Unlike other methods, MK-TGAN leverages prior knowledge graphs by exploiting graph neural networks. Our results show that prior knowledge integration strategies improve performance, and that MK-TGAN consistently produces synthetic samples with superior realism and biological plausibility.

Francesca Pia Panaccione, S. Mongardi, M. Masseroli et al. · 0 citations
Open access Aug 2025

Can synthetic data reproduce real-world findings in epidemiology? A replication study using adversarial random forests

Abstract Background Synthetic data hold substantial potential to address practical challenges in epidemiology due to restricted data access and privacy concerns. However, many current methods suffer from limited quality, high computational demands, and complexity for non-experts. Furthermore, common evaluation strategies for synthetic data often fail to directly reflect statistical utility and measure privacy risks sufficiently. Against this background, a critical underexplored question is whether synthetic data can reliably reproduce key findings from epidemiological research while preserving privacy. Methods We propose adversarial random forests (ARF) as an efficient and convenient method for synthesizing tabular epidemiological data. To evaluate its performance, we replicated statistical analyses from six epidemiological publications covering blood pressure, anthropometry, myocardial infarction, accelerometry, loneliness, and diabetes, from the German National Cohort (NAKO Gesundheitsstudie), the Bremen STEMI Registry U45 Study, and the Guelph Family Health Study. We further assessed how dataset dimensionality and variable complexity affect the quality of synthetic data, and contextualized ARF’s performance by comparison with commonly used tabular data synthesizers in terms of utility, privacy, generalization, and runtime. Results Across all replicated studies, results on ARF-generated synthetic data consistently aligned with original findings. Even for datasets with relatively low sample size-to-dimensionality ratios, replication outcomes closely matched the original results across descriptive and inferential analyses. Reduced dimensionality and variable complexity further enhanced synthesis quality. ARF demonstrated favourable performance regarding utility, privacy preservation, and generalization relative to other synthesizers and superior computational efficiency. Conclusions In summary, ARF reliably generates high-quality synthetic data that replicate diverse epidemiological analyses while offering a competitive privacy–utility trade-off.

J. Kapar, Kathrin Günther, L. Vallis et al. · 1 citation
Open access Jul 2026

Privacy-Preserving GAN for Synthetic Data against Membership Inference Attack

Experimental results demonstrate that PPGM-GAN outperforms state-of-the-art privacy-preserving generative models, producing high-utility synthetic data under the same privacy constraints.

Guizhang Cui, Guowei Wu, Lin Yao et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.