Explainable Predictive Maintenance for Smart Factories Using FPGA-GPU Edge–Cloud Hybrid Computing
Predictive maintenance in smart factories requires not only high prediction accuracy but also low-latency processing, adaptability to data drift, and understandable explanations for field operators. However, many existing approaches remain cloud-centered, label-dependent, and weak in practical explainability. This study proposes an explainable predictive maintenance framework based on FPGA-GPU Edge–Cloud hybrid computing for large-scale multivariate time-series environments. In the proposed system, FPGA modules perform streaming-oriented signal preprocessing and low-latency feature extraction, while GPU modules execute deep learning-based anomaly detection and fault prediction. To reduce dependence on labeled fault data, the framework incorporates masked autoencoder-based self-supervised representation learning. To improve long-term robustness in changing manufacturing environments, the framework also considers continual learning based on Elastic Weight Consolidation. In addition, a lightweight large language model with parameter-efficient fine-tuning and retrieval-augmented generation generates root-cause-oriented explanations and maintenance guidance. The method is organized as an integrated pipeline that combines data acquisition, edge preprocessing, temporal inference, explanation generation, and cloud-assisted model adaptation. The evaluation framework includes certification-oriented testing, comparative analysis, latency and throughput analysis, drift response analysis, and explainability assessment. According to a third-party test report, the proposed system achieved an event recall of 0.9822, event precision of 0.9529, event F1-score of 0.9674, and a false alarm rate of 0.000205. These results indicate that the proposed framework is practically feasible for real-time, explainable, and deployable predictive maintenance in smart factory environments.