Skip to content

When T2I Synthetic Data Backfires: Amplified Privacy Risks in Real-Synthetic Mix Training

Jul 2026 · arXiv.org · Vol abs/2607.13541 · 0 citations · 99 references
Computer Science

TL;DR

This work reveals, for the first time, that RSMT can substantially amplify privacy leakage of these real training samples, and proposes a lightweight leakage propensity indicator computable from real data alone that reliably identifies high-risk datasets unsuitable for entering RSMT, as a self-assessable mitigation.

Abstract

To overcome data scarcity and privacy constraints in data collection, it has become standard practice across academia and industry to augment real training data with text-to-image (T2I)-generated synthetic data, a paradigm we term Real-Synthetic Mix-Training (RSMT). While substituting synthetic data for sensitive real samples is widely regarded as a means to mitigate privacy exposure of the substituted data, the risk to the remaining real samples that actively participate in training has remained largely unexamined. This work reveals, for the first time, that RSMT can substantially amplify privacy leakage of these real training samples. We establish a theoretical framework, RSMT Memorization Amplification, proving that incorporating synthetic data displaces real samples toward peripheral regions of the mixed feature space, in turn forcing the model to memorize them more aggressively. Guided by this foundation, we propose RSMixLeak to systematically assess this risk through membership inference attacks (MIAs). RSMixLeak comprises two variants depending on the adversary's capability. The non-adversarial variant audits a benign RSMT pipeline with an honest T2I provider, establishing a lower bound on the leakage induced by the intrinsic gap between real and T2I-generated data. The adversarial variant considers an adversary who controls the T2I model or contributes crafted data to the T2I provider, and deliberately enlarges this distributional gap on a target class via either high-level semantic attribute binding or imperceptible pixel-level coating, further amplifying leakage on real training data while improving downstream model utility. Motivated by these findings, we further propose a lightweight leakage propensity indicator computable from real data alone that reliably identifies high-risk datasets unsuitable for entering RSMT, as a self-assessable mitigation.

View source

Similar papers

Open access Jul 2026

Privacy-Preserving GAN for Synthetic Data against Membership Inference Attack

Experimental results demonstrate that PPGM-GAN outperforms state-of-the-art privacy-preserving generative models, producing high-utility synthetic data under the same privacy constraints.

Guizhang Cui, Guowei Wu, Lin Yao et al. · 0 citations
Review Open access 2024

Synthetic Data Generation Strategies for Pipeline Robustness

A comprehensive survey and analysis of synthetic data generation techniques, classifying them by data modality, generation method, and application purpose and exploring how synthetic data contributes to the resilience of ML pipelines against failure modes such as concept drift, noise, and adversarial attacks.

Dennis M. Ritchie, Allen Newell · 0 citations
2026

[Comparing synthetic data generation methods in pharmacoepidemiology: reconciling reproducibility with privacy protection].

This study compares two methods of generating tabular synthetic data: synthpop, based on transparent and interpretable inferential methodologies, and Conditional Tabular-Generative Adversarial Networks (CT-GANs), which leverage deep learning approaches to reproduce complex multivariate distributions.

Flavia Mayer, Maria Laura Fazio, M. Cutillo et al. · 0 citations
Jul 2026

Optimal Domain-Aware Privacy Mechanisms for Synthetic Data Generation

This work considers normalized histograms as distribution estimators and characterize the asymptotically optimal domain-aware privacy mechanism within a specific class of DP mechanisms, and introduces PubMix, a public-data-aware DP mechanism that can be used in histogram-based data synthesis pipelines.

Sajani Vithana, Sangwon Jung, Haoyang Hu et al. · 0 citations
#artificial intelligence Preprint Aug 2026

Hidden Threat in Synthetic Data: Covert Targeted Bias Injection through Benign Text

This work constructs a pipeline in which a misaligned teacher model generates filtered synthetic datasets across domains such as creative writing and code generation, which are then used to fine-tune aligned student models, and shows that benign-looking synthetic data can act as a covert channel for transmitting targeted biases while largely preserving the student model's general task capabilities.

Minkyung Cho, Jihyo Kim, Seungwoo Song et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.