Aug 2026· Annual Review of Statistics and Its Application· 0 citations
TL;DR
This work discusses key considerations and methods for generating synthetic data, including sequential modeling, differentially private synthesis, deep generative models, and large language models, and outlines some open research challenges and future directions for synthetic data development.
Abstract
Synthetic data, i.e., data simulated from some statistical model, are an important tool for both privacy protection and artificial intelligence pipelines. In the privacy context, synthetic data enable agencies to disseminate record-level information while reducing disclosure risks. In the artificial intelligence context, synthetic data allow analysts to augment training sets, increase coverage of rare cases, and support experimentation when genuine data are scarce. For each usage, we discuss key considerations and methods for generating synthetic data, including sequential modeling, differentially private synthesis, deep generative models, and large language models. Throughout, we highlight key trade-offs between data usefulness, privacy protection, and model reliability. We conclude by outlining some open research challenges and future directions for synthetic data development.
The increasing demand for high-quality datasets for AI development has raised significant privacy, security, and regulatory concerns, particularly in sensitive domains such as healthcare, finance, and government. Synthetic data generation addresses these challenges by creating artificial datasets that preserve the statistical characteristics of real data while protecting individual privacy. Recent advances in generative AI, including GANs, VAEs, diffusion models, and transformer-based models, have significantly improved the realism and utility of synthetic data. This paper surveys synthetic data generation techniques, privacy-preserving methods, applications, research challenges, and future directions for developing trustworthy AI systems.
Per Brinch Hansen, Børge Diderichsen· International Journal of Mod...· 0 citations
This work considers normalized histograms as distribution estimators and characterize the asymptotically optimal domain-aware privacy mechanism within a specific class of DP mechanisms, and introduces PubMix, a public-data-aware DP mechanism that can be used in histogram-based data synthesis pipelines.
Sajani Vithana, Sangwon Jung, Haoyang Hu et al.· arXiv.org· 0 citations
A comprehensive survey and analysis of synthetic data generation techniques, classifying them by data modality, generation method, and application purpose and exploring how synthetic data contributes to the resilience of ML pipelines against failure modes such as concept drift, noise, and adversarial attacks.
Dennis M. Ritchie, Allen Newell· International Journal of Dat...· 0 citations
Synthetic data is seen as a promising solution for sharing data in sensitive contexts. However, recent work on privacy attacks have shown that there are still significant residual risks, especially for synthetic data generations methods that are not based on formal approaches such as differential privacy. In this paper, we investigate the privacy risks associated with local combination approaches for generating synthetic data in which synthetic profiles are built by combining real neighbouring profiles. More precisely, we focus on three methods from this family, namely SMOTE, Simulant and Avatar, which have been recently used as a way to share'anonymised data'in the healthcare domain. In particular, we conduct an extensive privacy analysis through a diverse set of attacks: membership inference, linkage and reconstruction attacks. Our results demonstrate substantial privacy leakage for all three methods, raising serious doubts about whether their outputs should be regarded as anonymous in practice.
H. Lautraite, Tristan Allard, Anne-Sophie Charest et al.· 0 citations
Synthetic data has become a common component of machine learning research. While widely adopted, its use in privacy-sensitive contexts has quietly shifted from a claim of residual inference risk under stated assumptions to an appearance-based property inferred from data generation itself. In this position paper, we argue that this shift reflects an implicit change in community standards for what counts as sufficient privacy evidence, rather than a misunderstanding of well-established privacy principles. Drawing on an empirical analysis of recent publications across major ML venues, we show that synthetic data is frequently used in privacy-sensitive settings without explicit articulation of threat models, inference risks, or falsifiable privacy claims. As a result, privacy assurance often remains implicit, difficult to verify, and unevenly distributed, with heightened exposure for rare and minority records. We argue for treating privacy as an explicit, evidence-based scientific claim and recommend that ML venues adopt norms requiring privacy-relevant assertions to be clearly scoped, testable, and contestable.
Neural networks are increasingly deployed in high-stakes applications with growing privacy leakage concerns. We show that this privacy leakage can occur even in the absence of representation imbalances that lead to traditional dataset biases. This poses significant privacy risks when deploying models that process sensitive attributes. In this context, we propose CutClean, a privacy-aware pruning method that allows to reduce privacy information flow through the network, while increasing its sparsity. Our approach employs auxiliary linear privacy heads placed at each network's block to quantify information leakage, and further applies increasing levels of sparsity to remove the private attribute leakage, measured in terms of the accuracy of the privacy head attached to the last block. Experiments on synthetic and real-world datasets demonstrate that our approach effectively minimizes private information flow while achieving high sparsity rates and preserving classification target accuracy.
Leonardo Magliolo, Vito Paolo Pastore, Giuseppe Valenzise et al.· International Conference on...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.