ChatRCA: A Root Cause Analysis Method via LLMs-based Multi-Agent with Human-in-the-Loop
Abstract
Root cause analysis (RCA) underpins cloud reliability by correlating anomalous signals to identify incident causes. Existing LLM-based RCA methods show promise, but monolithic LLM pipelines often under-capture RCA's structured workflow, while fully automated reasoning lacks calibrated uncertainty handling and expert oversight. We first conduct an empirical study to understand real RCA workflows, diagnostic roles, and useful forms of human feedback. Based on the findings, we propose ChatRCA, a multi-agent RCA framework that decomposes RCA into specialized subtasks handled by Manager, Observation, Architecture, Operation, and Expert agents. To mitigate uncertain or conflicting outputs, ChatRCA introduces lightweight human-in-the-loop feedback at two checkpoints: work-order verification and root-cause adjudication, using a consensus-then-arbitration protocol among three blinded operations engineers. We evaluate ChatRCA on three datasets, including TrainTicket, a private CMCC cloud-operation dataset, and GAIA. ChatRCA achieves strong performance in both service/component localization and root-cause category prediction, with category Top-1 accuracy of 91.11%, 86.67%, and 87.80% on the three datasets, respectively. On the CMCC dataset, ChatRCA also improves RCA explanation quality, achieving 63.45 BLEU-4 and 86.92 BERTScore. Deployment feedback further indicates its practical usefulness in cloud operations.