Skip to content

Author

Ketankumar Savajiyani

2 papers indexed here

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

Open access 2026

Agentic AI for Self-Healing Microservices: An LLM-Orchestrated Framework for Autonomous Fault Detection and Remediation in Regulated Enterprise Environments

: Enterprise-scale distributed microservices in regulated environments operate under stringent availability, latency, and auditability requirements. Traditional monitoring approaches detect anomalies reactively, after service degradation has already impacted end users or regulatory SLA (Service Level Agreement) obligations. This paper presents the Cognitive Agentic Self-Healing Platform (CASP): an LLM-orchestrated agentic system that continuously monitors distributed microservice environments, autonomously diagnoses fault conditions across six failure categories, and executes targeted remediation actions without human intervention. CASP integrates a structured four-phase autonomous loop comprising detection, diagnosis, remediation, and verification, with an LLM decision engine supporting both cloud-hosted and locally deployable backends, enabling deployment in air-gapped environments subject to data residency requirements. Evaluation on a controlled synthetic workload, calibrated against published benchmark profiles, demonstrates 88.3% overall diagnostic accuracy, an 83% reduction in mean recovery time compared to no self-healing, and a 53% reduction compared to static rule-based approaches. Post-recovery latency stabilizes within 7 minutes. The locally deployable LLM backend achieves p95 decision latency of 4.9 seconds, within the 5-second operational SLA threshold, making it viable for production deployment in regulated environments subject to data residency requirements.

Ketankumar Savajiyani · 0 citations
Conference Jul 2026

A Phase-Aware Reliability and Observability Framework for Enterprise Ingestion-To-Publish Data Pipelines

Multi-stage event-driven data pipelines underpin mission-critical enterprise workflows across regulated industries globally. Existing approaches such as Apache Kafka, Resilience4j, and OpenTelemetry address individual reliability and observability concerns in isolation, leaving engineering teams without a unified model for phase-specific fault containment. When failures occur, engineers spend hours correlating logs across dozens of services before identifying which processing stage is responsible. This paper presents PAROF, a Phase-Aware Reliability and Observability Framework that addresses this gap by decomposing the ingestion-to-publish pipeline into four explicitly bounded phases: Ingestion, Transformation, Persistence, and Publish. Each phase is governed by purpose-built reliability primitives and mandatory observability instrumentation that stamps every log, metric, and trace with a phase identifier. The framework was validated on a controlled, production-representative testbed using Apache Kafka, Spring Boot microservices, PostgreSQL, Toxiproxy for network fault injection, and k6 for synthetic load generation calibrated to published enterprise deployment benchmarks [1] [2]. Experiments across five failure scenarios demonstrated a 73% reduction in mean time to resolution (MTTR), a 91% cut in cross-phase incident propagation, complete elimination of data loss, and a 40% reduction in incident response time compared to a baseline pipeline without phase-aware primitives.

Ketankumar Savajiyani · 1 citation

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.