Synthetic Tabular Data Generation for Data Storage Systems
The generation of synthetic tabular data has emerged as a critical technique for addressing data scarcity, privacy concerns, and the need for augmented training datasets in machine learning applications. This study investigates the applicability and effectiveness of state-of-the-art generative models for synthesizing performance metrics of data storage systems. We utilize a performance dataset encompassing hard disk drive sequential storage configurations characterized by critical indicators including input/output operations per second and latency measurements. Our research methodology follows a systematic approach. First, we establish a baseline using the Synthetic Data Vault library to understand fundamental generative capabilities for tabular data. Subsequently, we implement and evaluate three advanced diffusion-based architectures and one generative adversarial network approach: TabSyn, TabDiff, TabDDPM, and CTGAN. The experimental framework encompasses comprehensive quality assessment through both visual inspection and quantitative metrics. Visual evaluation includes comparative analysis of input/output operations per second and latency distributions, marginal distributions of feature values, and correlation structure preservation between synthetic and real datasets. Quantitative assessment leverages detection scores, shape and trend similarity measures, and machine learning efficacy scores. The experimental results provide insights into the strengths and limitations of each generative approach when applied to storage system performance data exhibiting complex multi-modal distributions and intricate feature correlations.