On the Performance of Phase I Dispersion Control Charts under Machine Learning Imputed Data
Abstract
Phase I analysis is essential to understand the variability of the process and determine its stability. Providing Phase II control charts with poor parameters’ estimates leads to a weak performance. In case of incomplete Phase I data, the problem of missing values must be dealt with before estimating the process parameters. Researchers commonly rely on the traditional Mean Substitution (MS) and/or the Stochastic Regression (SRG) imputation methods, whilst the number of studies exploiting machine learning algorithms in the SPC field is rapidly increasing. Accordingly, in this study, we consider two common and powerful machine learning‐based imputation methods; which are the k‐Nearest Neighbors (kNN) and Support Vector Regression (SVR). We compare their effect with the two traditional methods, the MS and the SRG methods, on the performance of the G‐chart designed to monitor the process variability. Our results show that kNN imputation either surpasses the performance of the traditional methods or provides a similar performance. An application of the G‐chart is also illustrated. We recommend the use of the kNN imputation while monitoring the process dispersion.