Skip to content

Author

Chris D. Nugent

1 paper indexed here

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

Open access 2026

Recommendation: How to Use Synthetic Data in Machine Learning or Decision Support

Data serves as the foundation of contemporary artificial intelligence (AI) systems, yet ethical and practical constraints often limit the availability and usability of real-world datasets. Synthetic data (SD) has emerged as a valuable solution, enabling the development and training of AI models without compromising privacy standards or ethical guidelines. However, many challenges remain to address, from generating high-quality SD using various approaches to investigating the impacts of data training on machine learning (ML) models. This study examines the impact of balancing real and SD and provides some recommendations that researchers can further utilise to improve the ML model’s training process. Three datasets, Mobile Health (MHealth), High-Energy Physics Mass (HEPMass), and US Company Bankruptcy Prediction (UCBP), were pre-processed to ensure compatibility and used as the basis for SD generation using Conditional Tabular Generative Adversarial Networks (CTGAN) and Tabular Variational Autoencoders (TVAE). The study employed a hybrid data generation approach, splitting the training data into varying proportions of real and SD, with performance evaluated through 5-fold cross-validation on Decision Tree (DT), Gaussian Naive Bayes (GNB), and Linear Support Vector (L-SVM) ML models. The results indicate the optimal balance of real and SD for maximising model performance. Analysis of benchmarking results across three datasets shows that combining 30% real data with 70% CTGAN-generated synthetic data achieves the highest accuracy and overall model performance. In contrast, when using TVAE-generated data, a 20% real and 80% synthetic split is recommended to maintain similar performance. Additional experiments conducted on 10% reduced subsets showed that while the primary trends persisted under limited-data conditions, the optimal real-to-synthetic data ratio became more sensitive to the specific dataset. Statistical significance of the model performance differences was further confirmed using paired t-tests across all evaluated mixing ratios. This research analyses these results and considers the state of the art to recommend how further synthetic data can be useful with machine learning models and which approaches to adopt.

Majid Liaquat, Chris D. Nugent, Ian Cleland et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.