Anomaly detection in SKA monitor and control data using deep learning models
Abstract
In the complex and distributed environment of the SKA Observatory, predictive maintenance will play a crucial role in the upcoming production phase, with the aim of forecasting failures and malfunctions by analyzing data in order to identify early warning signs. Anomaly detection acts as the core engine for this process, identifying unusual patterns or deviations from normal operating. Within this context, lacking years of historical production data, a large amount of CSP (Central Signal Processor) monitor and control data is being collected during CI/CD pipeline jobs. The data generated during these tests is invaluable for establishing a valid operational baseline. Therefore, a range of deep learning models are trained to recognize the normal operational baseline of the observatory’s equipment. Once this baseline is established, the system can identify any data points or patterns that deviate significantly from the standard behavior like resource spikes, memory leaks, increased network I/O and so on. Deep Learning models are also suitable for identifying non-linear or subtle correlations which are almost impossible to be identified from a human operator given the huge amount of data and lots of different classes. The results of such AI-models are then summarized and compared following standard benchmark metrics of classifier systems. At the end of this process, the radio-telescope operators will have a comprehensive and innovative snapshot of the whole monitor and control infrastructure whereas anomaly behaviors have been detected, and thus the next phase of issues diagnosis and prognosis can be started in order to understand the source of the issue, predict the potential impact and plan proactive actions within the global process of predictive maintenance.