Skip to content

Optimal Domain-Aware Privacy Mechanisms for Synthetic Data Generation

Jul 2026 · arXiv.org · Vol abs/2607.18021 · 0 citations · 45 references
Computer Science Mathematics

TL;DR

This work considers normalized histograms as distribution estimators and characterize the asymptotically optimal domain-aware privacy mechanism within a specific class of DP mechanisms, and introduces PubMix, a public-data-aware DP mechanism that can be used in histogram-based data synthesis pipelines.

Abstract

Differential privacy (DP) imposes fundamental trade-offs between privacy and statistical fidelity in synthetic data generation. While access to public data has been shown to improve these trade-offs empirically, existing approaches use public data only indirectly, through pre-processing (e.g., using pre-trained generative models) or post-processing steps (e.g., matching target statistics estimated from public datasets), while relying on domain-agnostic DP mechanisms. In this work, we lay the theoretical framework to study the principled incorporation of public data into DP mechanisms themselves. We consider normalized histograms as distribution estimators and characterize the asymptotically optimal domain-aware privacy mechanism within a specific class of DP mechanisms. We introduce PubMix, a public-data-aware DP mechanism that can be used in histogram-based data synthesis pipelines. Our experiments demonstrate that PubMix significantly improves synthetic data generation quality compared to domain-agnostic privacy mechanisms.

View source

Similar papers

Conference Jul 2026

Embedding-Space Anonymization for Privacy-Preserving AI Systems

This paper studies embedding-space privacy as a representation-level learning problem. Rather than altering raw records directly, the proposed framework applies embeddingspace transformation to full-record representations through Gaussian perturbation and adversarial representation sanitization. The method is evaluated through ablation across utility metrics, linkage attacks, attribute-inference attacks, and membership-inference tests. The primary empirical evaluation uses a synthetic fusion recommendation benchmark built from MovieLens [1], [2] 32M behavior and Adult-derived demographics [3], while a secondary synthetic medical benchmark is used to examine cross-domain transferability under more constrained conditions. The strongest results appear in the recommendation experiments. Under grouped demographic privacy evaluation, the combined condition preserves recommendation utility with $N D C G {@} K=0.6312$ while reducing exact and entity linkage from 0.7090/0.7204 to 0.0001/0.0000. Sensitive-target attacker performance remains near the majority baseline, supporting the claim of empirical privacy improvement without visible ranking degradation in that benchmark. The healthcare experiments also demonstrate meaningful embedding transformation and linkage reduction, though the current benchmark remains datalimited and therefore less conclusive for utility-focused evaluation. Overall, the findings support the conclusion that embeddingspace transformation can preserve downstream utility while substantially reducing linkage risk and sensitive-information recoverability under explicit attacker evaluation. The findings support embedding-space transformation as a practical privacypreserving strategy for embedding-driven AI systems under explicit attacker evaluation.

D. Panagoulias, Evangelia-Aikaterini Tsichrintzi, E. Sakkopoulos · 0 citations
Open access Jul 2026

Research on Attention-Based Dynamic Differential Privacy Machine Learning Methods

Amid the rapid proliferation of big data and machine learning technologies, concerns around data privacy breaches have grown increasingly acute. Differential privacy offers a mathematically rigorous framework for privacy preservation and has found broad adoption in a range of machine learning settings. Still, conventional approaches tend to rely on static, uniform noise injection strategies—failing to account for the fact that different features and stages of training may have vastly dissimilar privacy sensitivities. As a consequence, privacy budgets are often used inefficiently, and model utility suffers noticeably. In response to these limitations, we introduce an attention-driven dynamic differential privacy mechanism that enables adaptive allocation of privacy budgets. Our design aims to uphold strong protection without sacrificing model performance as much as before. Specifically, we build a lightweight multi-head attention component that dynamically evaluates each feature’s relevance and the current phase of training, adjusting the intensity of injected noise on the fly. This component is designed for easy integration into existing deep learning pipelines without major structural modifications. We evaluate our method on the MNIST benchmark, and results show clear improvements in classification accuracy under identical privacy constraints. Further, we provide visual evidence through attention heatmaps and noise variation trajectories, which together offer interpretable support for the adaptive mechanism's effectiveness.

Haining Shang · 0 citations
Aug 2026

Synthetic Data: A Tool for Privacy Protection and Model Empowerment

This work discusses key considerations and methods for generating synthetic data, including sequential modeling, differentially private synthesis, deep generative models, and large language models, and outlines some open research challenges and future directions for synthetic data development.

Yinyihong Liu, Jerome P. Reiter · 0 citations
Book Open access Aug 2026

Efficient Privacy Auditing for Generative Model via Local Information

Diffusion models have become the dominant approach for text-to-image generation, but their ability to memorize training data raises increasing concerns about privacy leakage. Differential privacy (DP) is widely adopted to mitigate such privacy risks during model fine-tuning, yet the practical privacy leakage of differentially private diffusion models remains difficult to assess, especially in black-box settings where only generated images are observable. Existing auditing methods for diffusion models typically rely on membership inference attacks based on whole-image similarity between generated samples and target images. However, such approaches may underestimate privacy leakage when similarity between generated images and training samples is concentrated in localized image regions rather than at the whole-image level. In this work, we explore privacy leakage in diffusion models fine-tuned with differential privacy from a black-box perspective. We propose an empirical privacy assessment framework that leverages local image information, instead of treating images as indivisible wholes, to improve the distinguishability of privacy leakage signals. To improve privacy auditing efficiency and reduce sampling variance, we further leverage the image inpainting interface of diffusion models to perform region-focused auditing in a fully black-box setting. Extensive experiments across different fine-tuning and auditing settings demonstrate that our approach provides more reliable empirical assessments of privacy leakage than whole-image-based auditing methods.

Jingnan Xu, Leixia Wang, Xiaofeng Meng · 0 citations
Open access Jul 2026

Privacy-Aware Adaptive Differential Privacy for Semantic Retrieval: A Pii-Aware Dynamic Budget Allocation Framework

PADP is presented, a sensitivity-aware perturbation framework inspired by differential privacy principles, which provides a plug-and-play, middleware framework that can be easily integrated into enterprise RAG pipelines without requiring costly computations for LLM fine-tuning and reconstruction of vector indices.

Seçkin Mandaci, Yılmaz Vural, Ö. Turna · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.