2024· International Journal of Data Engineering and Intelligent Computing· 0 citations
TL;DR
A comprehensive survey and analysis of synthetic data generation techniques, classifying them by data modality, generation method, and application purpose and exploring how synthetic data contributes to the resilience of ML pipelines against failure modes such as concept drift, noise, and adversarial attacks.
Abstract
Machine learning pipelines are increasingly deployed in high-stakes domains where robustness against data inconsistencies, bias, and distributional shifts is crucial. However, real-world datasets often suffer from limitations such as data scarcity, class imbalance, and privacy constraints. Synthetic data generation has emerged as a promising strategy to overcome these challenges and enhance pipeline robustness. This paper presents a comprehensive survey and analysis of synthetic data generation techniques, classifying them by data modality, generation method, and application purpose. We explore how synthetic data contributes to the resilience of ML pipelines against failure modes such as concept drift, noise, and adversarial attacks. Through empirical case studies across domains, we demonstrate the practical benefits and limitations of integrating synthetic data into training and evaluation pipelines. Finally, we discuss ethical considerations and outline future research directions toward building more robust, fair, and privacy-preserving ML systems using synthetic data.
The increasing demand for high-quality datasets for AI development has raised significant privacy, security, and regulatory concerns, particularly in sensitive domains such as healthcare, finance, and government. Synthetic data generation addresses these challenges by creating artificial datasets that preserve the statistical characteristics of real data while protecting individual privacy. Recent advances in generative AI, including GANs, VAEs, diffusion models, and transformer-based models, have significantly improved the realism and utility of synthetic data. This paper surveys synthetic data generation techniques, privacy-preserving methods, applications, research challenges, and future directions for developing trustworthy AI systems.
Per Brinch Hansen, Børge Diderichsen· International Journal of Mod...· 0 citations
This work discusses key considerations and methods for generating synthetic data, including sequential modeling, differentially private synthesis, deep generative models, and large language models, and outlines some open research challenges and future directions for synthetic data development.
Yinyihong Liu, Jerome P. Reiter· Annual Review of Statistics...· 0 citations
As biomedical research increasingly relies on data-intensive tools, the quality and utility of datasets are critical. Challenges such as imbalances, biases, and ethical or legal constraints often limit access to high-quality data. Synthetic data generation can help overcome these limitations. Here, we present a comparative analysis of generative models for transcriptomic data, investigating strategies to incorporate prior biological knowledge via gene graphs. This ensures that synthetic data capture real-world gene patterns, maintaining their usefulness for downstream tasks. In particular, we introduce and benchmark three variants of the Generative Adversarial Network. Among the alternatives, MK-TGAN - an innovative multi-kernel, Graph Neural Network-based model - stands out for its performance in terms of both the realism and utility of the generated data. Unlike other methods, MK-TGAN leverages prior knowledge graphs by exploiting graph neural networks. Our results show that prior knowledge integration strategies improve performance, and that MK-TGAN consistently produces synthetic samples with superior realism and biological plausibility.
Francesca Pia Panaccione, S. Mongardi, M. Masseroli et al.· 0 citations
Experimental results demonstrate that PPGM-GAN outperforms state-of-the-art privacy-preserving generative models, producing high-utility synthetic data under the same privacy constraints.
Guizhang Cui, Guowei Wu, Lin Yao et al.· ACM Transactions on Privacy...· 0 citations
This work reveals, for the first time, that RSMT can substantially amplify privacy leakage of these real training samples, and proposes a lightweight leakage propensity indicator computable from real data alone that reliably identifies high-risk datasets unsuitable for entering RSMT, as a self-assessable mitigation.
Na Li, Boyu Kuang, Hongsheng Hu et al.· arXiv.org· 0 citations
Advances in diffusion-based generative models have motivated the use of synthetic image generation to alleviate data scarcity in vision tasks. While this strategy has shown promise in natural image benchmarks such as ImageNet, its effectiveness in sparse, high-variance real-world domains remains unclear. In this work, we focus on domains where images differ substantially from common image datasets and additional data are expensive to obtain. Against non-generative data augmentation baselines, we evaluate the downstream classifier performance improvements yielded by two schools of generative sparse data extension: distribution modeling and sample perturbation. Across five trauma classification tasks using subject-wise train--validation splits, no generative approach consistently outperforms a strong non-generative baseline. Feature-space analysis reveals recurring failure modes: memorization or collapse, distributional drift, and generation of visually plausible but simplified canonical instances that are easier to classify than real data.
Edward Zhang, Marcel Hussing, Tanay Tandon et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.