Skip to content

Designing Self-Healing AI Agentic Systems: A Framework for Autonomous Detection and Response

2026 · International Journal of Scientific Research and Management · Vol 14, pp. 2912-2919 · 1 citation

TL;DR

A new scientific object – the Autonomous Recovery Efficiency Score (ARES) – is introduced – a quantitative measure of autonomous resilience, as well as a supporting foundation for future autonomous self-healing AI agentic infrastructure.

Abstract

As AI agentic systems become more prevalent in distributed computing environments, they present unique challenges in terms of resilience, fault tolerance, and self-sustaining operations. In modern AI infrastructures, which are frequently subjected to cascading failures, communication disruptions, resource overload, and adaptive instability, particularly in the case of dynamic and large-scale environments. The existing methods in the recovery area are primarily fault detection and rule-based recovery, and provide little support to autonomous self-healing, adaptive recovery learning, and intelligent reconfiguration of the systems. In this work, the authors aim to fill in those missing pieces by proposing a new scientific paradigm, the Autonomous Cognitive Self-Healing Layer (ACSHL), which enables autonomous anomaly detection, cognitive fault diagnosis, orchestration of adaptive responses, and reinforcement-based recovery in AI agentic ecosystems. The proposed framework is developed under the Autonomous Resilient Agentic Intelligence (ARAI) theory, which intends to come up with a common model for resilient autonomous intelligence infrastructures. ACSHL brings together Intelligent Monitoring Agents, Cognitive Failure Analyzers, Adaptive Recovery Orchestrators, and Resilience Learning Engines, thereby allowing operations to adapt and self-recover continuously. The framework is then experimentally evaluated in the context of distributed fault injection environments, synthetic anomaly scenarios, Google cluster traces, and NASA system failure datasets. The performance of the systems was assessed using the metrics of detection accuracy, recovery latency, uptime stability, and efficient adaptive recovery. From experimental results, the proposed ACSHL framework has demonstrated recovery accuracy with 96.2% and less downtime by 41% as compared to conventional static and rule-based recovery systems and has also exhibited better adaptive recovery performance. The study introduces a new scientific object – the Autonomous Recovery Efficiency Score (ARES) – a quantitative measure of autonomous resilience, as well as a supporting foundation for future autonomous self-healing AI agentic infrastructure.

Read PDF

Similar papers

May 2026

A Self-Healing Framework for Reliable LLM-Based Autonomous Agents

Although large language model (LLM)-based autonomous agents are increasingly being utilized in complex software systems, ensuring their reliability remains a critical challenge due to unpredictable defects such as hallucinations, execution errors, and inconsistent reasoning. This study proposes a reliability-aware self...

Cheonsu Jeong, Young-Hyo Shin · 2 citations
Review Aug 2026

When Agentic AI Meets Integrated Sensing and Communication

Agentic artificial intelligence (AI) is transforming Integrated Sensing and Communication (ISAC) from a function-oriented physical-layer technology into a goal-driven, closed-loop intelligent system, a paradigm we term AISAC. Existing work on learning-based sensing, resource allocation, reconfigurable intelligent surfa...

Kai Li, Cong-Gai Li, S. A. Siddiqui et al. · 0 citations
Open access 2026

Agentic AI: Architectures, Types, Capabilities, Mathematical Equations and Governance in the Era of Autonomous Intelligence

A comprehensive framework for the design, evaluation, and responsible deployment of Agentic AI is proposed, emphasizing safety, explainability, human-in-the-loop supervision, and ethical compliance and aims to maximize the benefits of Agentic AI while minimizing potential risks.

Nitin S. Shrirao, Dnyaneshwar S. Jadhav, Sarita B. Patil · 0 citations
Review Open access Aug 2026

Decentralized Coordination Architectures for Intelligent Agent Swarms

The discussion treats distributed consensus, event-triggered communication, resilient control, fault-tolerant design, and cognition-inspired adaptation as parts of one architecture problem.

Tianwen Ge · 0 citations
Sep 2026

Toward Self-Evolving Agentic AI for ISAC-Enabled Low-Altitude Wireless Networks

This paper proposes a self-evolving agentic artificial intelligence (AI) framework for low-altitude wireless networks (LAWNs), introducing integrated sensing and communication (ISAC) into a unified self-evolution paradigm that transforms static foundation models with passive perception into fully autonomous, self-evolv...

Shi-Yi Gu, Lei Feng, Zhi-Xiang Yang et al. · 1 citation
Preprint Sep 2026

Agentic AI Enabling Autonomous, Self-Organizing, and Evolving UAV Networks

As low-altitude applications expand across emergency response, intelligent transportation, and autonomous operations, they demand communication networks that can deliver flexible, resilient, and rapidly deployable connectivity. Heterogeneous UAV networks are a promising solution, as they can dynamically provide sensing...

Zhao-Yang Li, Xin Jin, Zi-Jiu Yang et al. · 0 citations

Related blog posts

GPT-Lab Aug 28, 2026

We built an AI factory for HVAC control

What does it take to trust AI-driven HVAC optimization? Our AI Model Factory combines agents, machine learning, reinforcement learning and deterministic checks in a governed workflow designed for messy, real-world building data. The post We built an AI factory for HVAC control appeared first on GPT-Lab.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.