Aug 2026· International Journal for Research in Applied Science and Engineering Technology· 0 citations
TL;DR
The proposed Operational Decision-Readiness and Verification framework represents a deployment claim through analytical capability, operational authority, and assurance maturity and is conceptual and requires prospective and inter-rater validation.
Abstract
Artificial intelligence for information technology operations has progressed from alert correlation and anomaly
detection to root-cause analysis, mitigation recommendation, and agents that can invoke infrastructure tools. This progression
creates an assurance problem: analytical performance does not establish authority to change a live cloud system. This critical
narrative review examines the evidence required before AI output may influence or execute failure-management action in
mission-critical cloud infrastructure. Searches through 24 July 2026 covered scholarly databases, major systems venues,
standards sources, and authoritative production reports. Forty-four sources were coded by operational task, evidence setting,
authority, safety control, rollback, and outcome verification. Production evidence is substantial for detection, triage, diagnosis,
and several narrowly bounded mitigation systems, but remains weak for general-purpose autonomous action. Recent agentic
AIOps surveys emphasize contracts, bounded tools, canary deployment, and rollback; the unresolved need is an assessment
method that separates what a system can infer, what it may do, and what evidence shows that the action remained safe. The
proposed Operational Decision-Readiness and Verification framework represents a deployment claim through analytical
capability, operational authority, and assurance maturity. Twelve domains use explicit 0-3 evidence anchors, while endpoint
validity, evidence integrity, intervention risk, authorization, reversibility, and recovery verification operate as non-compensatory
gates. Worked assessments of six systems illustrate distinctions among analytical support, testbed execution, and narrow
production autonomy. The framework is conceptual and requires prospective and inter-rater validation; it structures an actionspecific readiness case rather than certifying a model or product
Distributed cloud infrastructure managing workloads across multi-datacenter environments at enterprise scale encounters failure modes, node degradation, network partition, resource contention, storage latency spikes, that reactive fault management detects only after service impact has occurred. At hundred-thousand-host...
Arjun Danda Sureshbabu· International journal of com...· 0 citations
Security Operations Centers (SOCs) are essential for monitoring and responding to cyber threats in cloud-native environments, where infrastructure is dynamic, multi-tenant, and API-driven. Conventional SOCs rely heavily on manual triage and SIEM-based alerting, resulting in delayed detection of cloud-specific attacks s...
Jilika Jithendarnadh, J. Balaraju· Review of Computer Engineeri...· 0 citations
Distributed enterprise technology environments increasingly depend on interconnected applications, cloud services, APIs, databases, infrastructure platforms, and automated deployment pipelines to sustain continuous digital operations. While this architecture improves scalability and agility, it also increases exposure...
Bibane Jaree· International Journal of Res...· 0 citations
The study supports the use of bounded, auditable agentic control for recurring operational failure modes, and suggests that policy-gated autonomy lowers mean time to resolve (MTTR) and incident recurrence.
VenkateswaraReddy Gudise· American Journal of Technolo...· 0 citations
AI agents can diagnose cloud incidents, synthesize operational commands, and invoke state-changing APIs, but a plausible remediation is not necessarily safe to execute. This study presents RACER, a runtime-assurance mechanism that treats every AI-generated repair as an untrusted proposal until it is bound to a machine-...
Prudvi Saisaran Ponduru, Pavani Priya Vyshnavi Nandanavanam, Sai Kesav Kumar Ponduru· International Journal of Adv...· 0 citations
Security Operations Centers (SOCs) increasingly rely on Security Orchestration, Automation, and Response (SOAR) platforms to manage high-volume alerts, enrich telemetry, execute playbooks, and shorten incident-response cycles. However, many deployed SOAR systems remain rule dominated: actions are triggered by static if...
Ikenna Mbuko, O. Ijiga, L. Enyejo· International Journal of Eng...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.