This work reveals, for the first time, that RSMT can substantially amplify privacy leakage of these real training samples, and proposes a lightweight leakage propensity indicator computable from real data alone that reliably identifies high-risk datasets unsuitable for entering RSMT, as a self-assessable mitigation.
Abstract
To overcome data scarcity and privacy constraints in data collection, it has become standard practice across academia and industry to augment real training data with text-to-image (T2I)-generated synthetic data, a paradigm we term Real-Synthetic Mix-Training (RSMT). While substituting synthetic data for sensitive real samples is widely regarded as a means to mitigate privacy exposure of the substituted data, the risk to the remaining real samples that actively participate in training has remained largely unexamined. This work reveals, for the first time, that RSMT can substantially amplify privacy leakage of these real training samples. We establish a theoretical framework, RSMT Memorization Amplification, proving that incorporating synthetic data displaces real samples toward peripheral regions of the mixed feature space, in turn forcing the model to memorize them more aggressively. Guided by this foundation, we propose RSMixLeak to systematically assess this risk through membership inference attacks (MIAs). RSMixLeak comprises two variants depending on the adversary's capability. The non-adversarial variant audits a benign RSMT pipeline with an honest T2I provider, establishing a lower bound on the leakage induced by the intrinsic gap between real and T2I-generated data. The adversarial variant considers an adversary who controls the T2I model or contributes crafted data to the T2I provider, and deliberately enlarges this distributional gap on a target class via either high-level semantic attribute binding or imperceptible pixel-level coating, further amplifying leakage on real training data while improving downstream model utility. Motivated by these findings, we further propose a lightweight leakage propensity indicator computable from real data alone that reliably identifies high-risk datasets unsuitable for entering RSMT, as a self-assessable mitigation.
Experimental results demonstrate that PPGM-GAN outperforms state-of-the-art privacy-preserving generative models, producing high-utility synthetic data under the same privacy constraints.
Guizhang Cui, Guowei Wu, Lin Yao et al.· ACM Transactions on Privacy...· 0 citations
A comprehensive survey and analysis of synthetic data generation techniques, classifying them by data modality, generation method, and application purpose and exploring how synthetic data contributes to the resilience of ML pipelines against failure modes such as concept drift, noise, and adversarial attacks.
Dennis M. Ritchie, Allen Newell· International Journal of Dat...· 0 citations
This study compares two methods of generating tabular synthetic data: synthpop, based on transparent and interpretable inferential methodologies, and Conditional Tabular-Generative Adversarial Networks (CT-GANs), which leverage deep learning approaches to reproduce complex multivariate distributions.
Flavia Mayer, Maria Laura Fazio, M. Cutillo et al.· Epidemiologia & Prevenzione· 0 citations
This work considers normalized histograms as distribution estimators and characterize the asymptotically optimal domain-aware privacy mechanism within a specific class of DP mechanisms, and introduces PubMix, a public-data-aware DP mechanism that can be used in histogram-based data synthesis pipelines.
Sajani Vithana, Sangwon Jung, Haoyang Hu et al.· arXiv.org· 0 citations
Anti Adversarial Training (AT-AT), a training regime that intentionally learns non-robust features to obtain both superior reconstruction defense and higher accuracy than state-of-the-art defenses, is introduced.
Rasmus Torp, Shailen K. Smith, Adam Breuer· arXiv.org· 0 citations
This work constructs a pipeline in which a misaligned teacher model generates filtered synthetic datasets across domains such as creative writing and code generation, which are then used to fine-tune aligned student models, and shows that benign-looking synthetic data can act as a covert channel for transmitting targeted biases while largely preserving the student model's general task capabilities.
Minkyung Cho, Jihyo Kim, Seungwoo Song et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.