2026· E3S Web of Conferences· 0 citations· 11 references
TL;DR
A deep learning-based failure prediction model that integrates Convolutional Neural Networks and Bidirectional Long Short-Term Memory networks to identify job failures before they occur is presented, improving the performance of cloud computing applications by reducing job failures and optimising resource utilisation.
Abstract
Cloud service providers face significant challenges in preventing hardware and software failures due to the large-scale and heterogeneous nature of cloud computing. Although many studies have focused on characterising failed jobs, fewer have explored proactive failure prediction. This paper presents a deep learning-based failure prediction model that integrates Convolutional Neural Networks (CNN) and Bidirectional Long Short-Term Memory (BiLSTM) networks to identify job failures before they occur. The proposed model improves the performance of cloud computing applications by reducing job failures and optimising resource utilisation. Using the Google Cluster Traces dataset, we analyse failure patterns and evaluate the effectiveness of the model across multiple performance metrics. The results demonstrate the robustness of the proposed scheme, achieving an accuracy of 99.96%, along with a high F1-score of 99.92% when compared to existing models. These findings highlight the potential of deep learning in proactive failure mitigation, providing a foundation for future advances in cloud workload reliability.
A dynamic recurrent neural network is proposed to accurately predict workloads and integrates an auto-encoder to effectively extract representations from the original workload data with high dimensionality to enable adaptive and accurate predictions for highly variable workloads.
Okore Kalu, C. Okafor, P. Asuquo et al.· E3S Web of Conferences· 0 citations
A deep learning-based model for task scheduling in cloud computing that employs a convolutional neural network to predict the optimal machines for task allocation and consumes less energy than other models is proposed, demonstrating its effectiveness in cloud task scheduling.
Kavita Rani, O. Sangwan, R. Garg· IAES International Journal o...· 0 citations
Cloud infrastructure plays a pivotal role in delivering uninterrupted computing services across the globe. However, unexpected failures in critical components such as cooling systems, networking hardware, and power supply units can result in significant downtime and economic loss. This paper investigates the application of Artificial Intelligence (AI) and Machine Learning (ML) techniques for predictive maintenance in cloud data centers. By analyzing historical sensor data, system logs, and performance metrics, various ML models—ranging from classical approaches like Random Forests to deep learning models such as LSTM and Autoencoders—are deployed to predict potential failures and trigger maintenance alerts. The paper also explores challenges like data imbalance, real-time inference, and deployment at scale in cloud environments. Experimental evaluations demonstrate the efficacy of ML models in identifying fault patterns, thereby improving operational reliability, reducing maintenance costs, and minimizing service disruptions.
Rakesh T· International Journal of Art...· 0 citations
Accurate workload prediction in cloud data centers is essential for efficient resource management, yet high-dimensional and noisy operational data often hinder forecasting performance. This work extends the original CVCBM model by integrating a lightweight Bidirectional GRU (BiGRU) with Bidirectional LSTM (BiLSTM) to enhance prediction efficiency while maintaining temporal feature extraction. Initially, workload signals are denoised and decomposed using a two-stage process—Complete Ensemble Empirical Mode Decomposition with Adaptive Noise (CEEMDAN) followed by Variational Mode Decomposition (VMD). Sample Entropy (SE) selects meaningful components, and K-Means clustering prioritizes high workload data for training. The hybrid Conv1D-BiLSTM-BiGRU architecture captures multi-scale temporal patterns and both short-term and long-term dependencies. The trained model is deployed using the Flask framework for real-time workload prediction, allowing interactive input of datasets and immediate forecasting. Experimental evaluation demonstrates that the extended model reduces computational overhead while improving prediction accuracy, providing robust, scalable, and real-time forecasting for cloud data center resource management.
Rayala Ashok, M. Praveena, G. S. Prasad et al.· 2026 7th International Confe...· 0 citations
With the explosive development of LLM-empowered agent technology, LLM inference performance has become more important than training. Cloud computing is a popular deployment approach, where performance prediction is vital for instance selection and QoS assurance. However, prediction is challenging due to GPU hardware heterogeneity, Transformer operator variations, and dynamic inference configurations. Virtualization and other features vary across clouds, further increasing prediction difficulty. Existing methods suffer from low accuracy and poor generalization. To tackle these issues, we propose Dispeller, a prediction model for GPU-accelerated cloud environments with three feature sets: 1) basic GPU hardware feature with 7 dimensions; 2) operator-level GPU performance feature with 4 dimensions; 3) inference configuration feature with 4 dimensions. We conduct experiments on public cloud GPUs and collect a real-world dataset of 10,112 samples. Random Forest is adopted to learn the nonlinear mapping between features and performance. Experimental results show that Dispeller achieves high prediction accuracy with TPS $\mathrm{R}^{{2}} = 0.951$ on seen GPUs and 0.989 on unseen GPUs, demonstrating strong cross-GPU generalization. An ablation study confirms that inference configuration features contribute 87.3% of the predictive power. Dispeller is therefore able to recommend cloud resources and optimize LLM deployment costs.
Huan Zhou, Zhi-Peng Wang, Meng-Juan Li et al.· Fall Joint Computer Conferen...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.