Author

Dhananjay M. Kanade

1 paper indexed here

Fetches their full publication history.

Not the right person? Other researchers publish under this name.

Open access Jun 2026

Decision Tree-Based Synthetic Data Generation Framework for Privacy-Preserving Data Publishing

The growing demand for data sharing in healthcare, scientific research, and public policy raises an ongoing challenge of protecting individual privacy while keeping data useful for analysis. This research introduces a Decision Tree-based synthetic data generation method, DTSDG and evaluates its performance using six well-known anonymization techniques. , namely, ‘k-Anonymity’, ‘l-Diversity’, ‘t-Closeness’, ‘Differential Privacy (DP)’,  ‘Generative Adversarial Networks (GANs)’, and ‘Copula-GAN’ The proposed method creates synthetic records step by step. Each attribute is predicted using a trained decision tree, which offers a balance of interpretability, efficiency, and accuracy. A composite scoring framework combines normalized Nearest Neighbor Distance (NND), Jensen–Shannon Divergence (JSD), and Wasserstein Distance into a unified metric is employed to assess how effectively each of these methods balance the conflicting performance measures: the privacy preservation and the data utility. Experiments were conducted using six diverse datasets, namely, ‘Adult’, ‘ATUS’, ‘FARS’, ‘Breast Cancer Wisconsin’, ‘cardiovascular disease’, ‘Heart Disease Cleveland’. The experimental results demonstrate that the proposed DTSDG method consistently achieves stable and competitive scores across all these datasets providing the best balance between privacy preservation and the data utility. The proposed DTSDG approach could be a practical and promising solution for publishing data while protecting privacy. This is especially true in situations where computational resources are limited or there is a strong need for model transparency.

Dhananjay M. Kanade, Dr. Shirish S. Sane, Dr. Uday Wad · 0 citations