Skip to content

TriCalRAG: A Three-Strategy, Retrieval-Augmented Benchmark for On-Premise LLM-Based Root Cause Analysis in AIOps

Sep 2026 · 0 citations · 50 references
Computer Science

TL;DR

This work presents TriCalRAG, a benchmark evaluating open-weight LLMs served locally via vLLM on a single high-memory workstation GPU against a classical LSTM-based log anomaly detector (DeepLog), across four real, publicly available log datasets (BGL, HDFS, Thun-derbird, OpenStack).

Abstract

Cloud-hosted large language models (LLMs) are increasingly used for root cause analysis (RCA) in AIOps pipelines, but they introduce data privacy risk, network latency, and per-query cost that scale poorly with production log volumes. We present TriCalRAG, a benchmark evaluating open-weight LLMs served locally via vLLM on a single high-memory workstation GPU (NVIDIA RTX PRO 6000, 96GB) against a classical LSTM-based log anomaly detector (DeepLog), across four real, publicly available log datasets (BGL, HDFS, Thunderbird, OpenStack). We evaluate two open-weight models (Qwen2.5-14B, Mistral-Small) under three prompting strategies: zero-shot, few-shot, and retrieval-augmented generation (RAG) over a labeled incident history, reporting accuracy, precision/recall, and F1 with bootstrap 95% confidence intervals across 3 random seeds, alongside throughput and VRAM footprint. Our results show that RAG not only improves mean F1 by 0.10-0.27 over zero-shot prompting but, more importantly, substantially stabilizes model calibration: zero-shot prompting drives both models toward near-degenerate behavior (predicting"anomaly"on up to 100% of incidents on some datasets), while RAG keeps predicted-positive rates close to the true class balance in the majority of configurations. Mistral-Small achieves higher macro-averaged F1 than Qwen2.5-14B (0.644 vs. 0.560) but exhibits calibration failures in more configurations (7 vs. 5 of 12), while running at roughly half the throughput - indicating the better model choice depends on whether a deployment prioritizes peak accuracy or predictable behavior across prompting conditions. Ablations show batching scales throughput 41 times on a single card and that 4-bit quantization reduces latency 20% with no measurable accuracy loss. We release our benchmark harness, dataset splits, and evaluation code to support reproducible on-premise AIOps research.

View source

Similar papers

Preprint Sep 2026

TriCalRAG: A Three-Strategy, Retrieval-Augmented Benchmark for On-Premise LLM-Based Root Cause Analysis in AIOps

Operational logs create a need for private, resource-efficient incident analysis, but aggregate detection scores can conceal severe prediction bias. We present TriCalRAG, a reproducible benchmark for log-anomaly detection with generated root-cause and remediation outputs across BGL, HDFS, Thunderbird, and OpenStack. Th...

Rohit Patel, S. K. Mohanty, Jeenal Chaudhary · 1 citation
#natural language process... Preprint Sep 2026

Beyond Fluent Generation: A CPU Reliability Benchmark for MCP-Style Tool Calling in Sub-2B Small Language Models for Edge Deployment

Resource constrained single-board computers including Raspberry Pi, NVIDIA Jetson Nano, Arduino UNO Q, Orange Pi, and LattePanda motivate on-device small language model (SLM) agents that reduce cloud dependence, improve data locality, and tolerate intermittent connectivity. Model Context Protocol (MCP)-style tool invoc...

Abrar Shahriar Qurat-Ul-Ain Mastoi · 0 citations
Review Open access Sep 2026

Energy-Efficient Retrieval-Augmented Generation: A Systematic Review of Retrieval, Reranking, Caching, and Context-Management Strategies

Retrieval-Augmented Generation (RAG) has become a standard paradigm for grounding large language models (LLMs) in external knowledge, but typical deployments introduce substantial energy, latency, and cost overhead due to expensive retrieval and context-processing pipelines. Recent work in "Green AI" and sustainable ma...

Anupam Dhakal, Prashant Pokharel, S. Adhikari · 0 citations
Open access Sep 2026

Raw CVE/CWE Retrieval Does Not Improve LLM-Based Vulnerability Detection in Python: A Pre-Specified Null Result and a Retrieval-Dose Audit

Retrieval-Augmented Generation (RAG) over vulnerability databases is widely expected to improve LLM-based vulnerability detection. We report a pre-specified evaluation in which it does not, and derive from it an evaluation protocol. On a near-balanced benchmark of 100 Python 3.12 snippets, raw CVE/CWE retrieval never m...

Patrick Deininger, S. Rappl, Helmut Lindner et al. · 0 citations
#machine learning Preprint Sep 2026

RAILS: Retrieval-Augmented Incremental LLM Clustering at Scale

RAILS is presented, a retrieval-augmented incremental LLM clusterer that turns clustering into a simple loop over a growing label pool and scales through document batching with bounded concurrency.

Armin Oliya, Aleksandra Sawczuk, Radosław Białobrzeski · 0 citations
#small language model Preprint Sep 2026

SSD-LLaMA: SSD-Native Inference for Trillion-Parameter MoE at 1+ Token/s on a Consumer PC

An SSD-native local MoE inference system that addresses challenges with an SSD I/O pipeline optimized for expert delivery, a native three-tier storage hierarchy that delivers and retains experts dynamically, and balanced CPU--GPU hybrid execution.

Fang-Zhou Liang, Yibin Shen, Jian-Min Hu et al. · 0 citations

Related blog posts

MIT News · Artificial Intelligence Sep 29, 2026

Who we become when we talk to machines

Professor Sherry Turkle’s new book, “Artificial Intimacy,” offers a withering critique of chatbots and the antisocial dynamics she believes they encourage.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.