Skip to content
Open access

AI-Powered Anomaly Detection in Cloud-Based Applications

2024 · International Journal of Scientific Research and Management · Vol 12, pp. 1102-1114 · 0 citations

TL;DR

The findings suggest that AI-powered anomaly detection significantly strengthens observability and security in cloud-based applications, enabling proactive threat mitigation and operational optimization in increasingly complex distributed environments.

Abstract

The rapid adoption of cloud-based architectures has increased system scalability and flexibility while simultaneously expanding the attack surface and operational complexity of modern applications. Traditional rule-based monitoring systems, which depend on static thresholds and predefined signatures, struggle to detect sophisticated threats and performance irregularities in highly dynamic, elastic, and ephemeral cloud environments where workloads scale up and down continuously and services are frequently redeployed. This paper explores the design and implementation of AI-powered anomaly detection frameworks tailored for cloud-native infrastructures, examining the theoretical foundations, architectural components, and practical deployment considerations of intelligent monitoring systems. By leveraging machine learning techniques such as supervised learning, unsupervised clustering, and deep learning models including recurrent neural networks and autoencoders, artificial intelligence systems can identify deviations from baseline behavior across distributed services, containers, and microservices in real time. The proposed approach integrates telemetry data from logs, metrics, and network traces to establish adaptive behavioral profiles, emphasizing automated feature extraction, continuous model training, and feedback loops that reduce false positives while improving detection accuracy. The framework is explicitly designed to operate across the full lifecycle of anomaly management, from raw data ingestion through model inference to alert generation and remediation. Additionally, this study addresses scalability challenges, data privacy considerations, and integration with DevOps and SecOps workflows. Experimental evaluation, conducted on a dataset exceeding 500,000 records drawn from logs, metrics, and network traffic, demonstrates improved detection rates, faster incident response times, and enhanced system resilience compared to conventional monitoring tools. Five model families were benchmarked side by side, with hybrid ensemble approaches achieving the strongest overall results. The findings suggest that AI-powered anomaly detection significantly strengthens observability and security in cloud-based applications, enabling proactive threat mitigation and operational optimization in increasingly complex distributed environments.

Read PDF

Similar papers

Open access Aug 2026

AI-Powered Scalable Anomaly Detection Framework for Secure Data Processing in Modern Cloud Architectures

An AI-powered scalable anomaly detection framework for secure data processing in cloud architectures that uses machine learning techniques such as Isolation Forest, Random Forest, and deep learning models to detect abnormal patterns in cloud data.

H. S, R. .Suganeswaran, S. M et al. · 0 citations
Jul 2026

ClouDens: Operational Context-Aware Anomaly Detection for Large-scale Cloud System Monitoring

With the rapid growth of cloud computing infrastructures in scale and complexity, network monitoring for Large-scale Cloud Systems (LCSs) has become increasingly challenging, requiring automated and reliable anomaly detection to maintain service availability. Modern LCSs continuously generate telemetry logs from distributed cloud services, producing high-dimensional multivariate time series that capture system operations. Detecting anomalies in this context is difficult due to extreme dimensionality, complex dependencies among distributed components, and severe sparsity from intermittently active services. Taking these challenges into account, we first conduct an empirical study on telemetry logs from the IBM Cloud Console platform, and then propose ClouDens, an anomaly detection framework tailored to LCS monitoring that leverages operational-context attributes encoded in the telemetry log schema to improve detection accuracy and early identification of anomalies. ClouDens partitions high-dimensional telemetry logs into domain-guided subsets, constructs a context-aware graph modeling operational service dependencies, and employs Spatio-Temporal Graph Neural Networks for forecasting-based anomaly detection. We evaluate ClouDens on the recently released IBM Cloud Telemetry Dataset and provide practical insights into designing reliable anomaly detection solutions for LCS monitoring. Results show ClouDens achieves higher NAB scores in count-based telemetry features, indicating more accurate, earlier anomaly detection with broader coverage than a GRU-based model. Our study further reveals that telemetry feature subsets, operational-context modeling, scoring strategies, and sparsity imputation all substantially influence detection performance, offering practical guidance for designing and fairly benchmarking anomaly detection approaches for LCS monitoring.

Thu T. H. Doan, Mohammad Saiful Islam, Andriy V. Miranskyy et al. · 0 citations
Open access Aug 2026

An AI-Based Framework for Predictive Scaling and Anomaly Detection in Enterprise Data Platforms

Enterprise data platforms face escalating challenges managing dynamic workloads, ensuring optimal resource allocation, and detecting anomalous behaviors that threaten system integrity and performance. Traditional reactive scaling approaches and rule-based anomaly detection systems struggle to cope with the complexity, velocity, and unpredictability of modern data processing environments. This research presents a comprehensive AI-based framework that integrates predictive scaling mechanisms with intelligent anomaly detection to optimize enterprise data platform operations. The framework employs machine learning algorithms including time-series forecasting models for workload prediction, reinforcement learning for dynamic resource allocation, and unsupervised learning techniques for anomaly identification. Through implementation and evaluation across five enterprise organizations managing data platforms processing over 2.8 petabytes daily, the framework demonstrated 67% reduction in resource over-provisioning costs, 43% improvement in query performance through proactive scaling, and 89% accuracy in detecting anomalous system behaviors with average detection latency under 45 seconds. The predictive scaling component accurately forecasted workload spikes 20-45 minutes in advance, enabling preemptive resource allocation that prevented performance degradation during demand surges. Anomaly detection modules identified security threats, data quality issues, infrastructure failures, and performance bottlenecks with significantly lower false positive rates than traditional threshold-based systems. However, implementation challenges emerged including model training data requirements, computational overhead of real-time predictions, integration complexity with legacy systems, and calibration needs across different workload patterns. This research contributes practical architectural designs, algorithm selection guidance, and operational frameworks for organizations seeking to implement AI-driven optimization in their data platforms while maintaining reliability and cost-effectiveness.

Sangeetha Mandapaka · 0 citations
Open access Jul 2026

Adaptive intrusion detection system for cloud security using deep learning

The findings confirm that the proposed IDSaaS framework provides an efficient, scalable, and adaptive solution for real-time cloud intrusion detection and significantly enhances the reliability and resilience of modern cloud and industrial cybersecurity infrastructures.

Unik B. Lokhande, Kavita Sonawane · 0 citations
Review Open access 2026

An Adaptive Machine Learning Framework for Anomaly Detection in Cloud-Based Storage Systems

Critical anomaly detection difficulties, such as false alarms during workload variations and delayed breach detection, have been brought about by the quick adoption of cloud-based storage systems. Data integrity and operational effectiveness are jeopardized by traditional static models' inability to adjust to the dynamic nature of cloud settings. In order to improve anomaly detection accuracy and resource optimization, this study created and verified an adaptive machine learning framework that makes use of real-time model updates and domain-specific cloud infrastructure information. CloudSim simulations of 1,000 cloudlets (10 runs, σ = 0.000), a quantitative survey of 51 IT specialists (92.7% response rate), and qualitative interviews with 13 infrastructure administrators were all included in the mixed-methods sequential explanatory design. A substantial importance-implementation gap in domain knowledge was found (Δ = 1.45, p <.001). With only 15% CPU overhead, the suggested framework, which is based on a domain-enhanced Random Forest, improved the F1-score by 64% and decreased false positives by 54% when compared to static thresholds. By bridging the gap between theoretical machine learning and the realities of cloud infrastructure in Africa, the study offers a deployable framework for improving cloud security in Kenya and other resource-constrained environments.

Milimo Moses Sibilike · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.