The Role of Synthetic Data in Educational Research: A Systematic Review
Abstract
The primary aim of this study is to provide a comprehensive and structured synthesis of existing research to understand how synthetic data is conceptualized, generated, and utilized within educational contexts. By analyzing 29 peer-reviewed articles, the research identifies seven primary dimensions of application: privacy and data sharing, data augmentation, NLP/ text generation, predictive modeling, pedagogical design, methodological analysis, and synthetic data in mobile, interactive, and adaptive learning systems. A significant finding is the increasing integration of artificial intelligence (AI) and machine learning technologies, such as generative adversarial networks (GANs) and large language models (LLMs), which are now central to generating high-fidelity artificial records and augmenting qualitative datasets. Across these analytical, predictive, and pedagogical domains, synthetic data offers a viable response to persistent challenges related to data scarcity, privacy constraints, and limited data accessibility in education. The findings indicate a growing reliance on synthetic generation as an emerging methodological response to data-intensive demands. While synthetic data supports advanced modeling, adaptive learning systems, and instructional design, its epistemological legitimacy and methodological robustness remain contingent on rigorous validation practices. The study concludes that the field currently lacks standardized validation protocols, particularly regarding subgroup equity and fairness. Establishing transparent, equity-aware frameworks remains essential for the future integration of synthetic data into applied educational systems.