Skip to content
Book Open access

LLM Agents for AIOps in Kubernetes: An Industrial Experience Report with Red Hat OpenShift

Jul 2026 · SIGSOFT FSE Companion · pp. 678-689 · 0 citations · 48 references
Computer Science

TL;DR

An industry experience report conducted on Red Hat OpenShift to explore how LLMs can be operationalized in real-world Kubernetes-based environments and integrates predictive machine learning models with LLM agents through tool-augmented reasoning, highlighting novel methods to automate IT tasks, enhance observability, and reduce operator burden.

Abstract

The integration of Artificial Intelligence (AI) into IT Operations Management (ITOM), commonly referred to as AIOps, offers substantial potential for automating workflows, enhancing efficiency, and supporting informed decision-making. However, practical implementation of AI within IT operations remains challenging, particularly due to data quality issues, the complexity of cloud-native environments, and skill gaps within operational teams. The emergence of Large Language Models (LLMs) presents new opportunities to address these barriers by leveraging their advanced natural language understanding, enabling the analysis of unstructured data such as logs, incident reports, and technical documentation. In this paper, we present an industry experience report conducted on Red Hat OpenShift to explore how LLMs can be operationalized in real-world Kubernetes-based environments. We integrate predictive machine learning models with LLM agents through tool-augmented reasoning, highlighting novel methods to automate IT tasks, enhance observability, and reduce operator burden. Our findings provide insights into both the capabilities and limitations of LLMs in production-grade AIOps scenarios.

Read PDF

Similar papers

Jul 2026

DataClawEval: A Benchmark for Data Engineering Agents in Real Industrial Harness

DataClawEval is introduced, the first comprehensive benchmark designed specifically to evaluate the end-to-end task completion capabilities of autonomous agents in real-world data engineering scenarios, and it comprises 100 rigorous, end-to-end tasks spanning five execution engines.

Debin Meng, Jiaming Yang, Zefang Zong et al. · 0 citations
Review Open access 2019

Explainable AI (XAI) Models for Transparent Decision Making in IIoT

The integration of Artificial Intelligence (AI) into the Industrial Internet of Things (IIoT) has enabled predictive analytics, autonomous control, and optimized operations. However, the increasing reliance on complex and opaque machine learning models raises concerns regarding trust, accountability, and regulatory compliance in critical industrial environments. Explainable AI (XAI) aims to address these concerns by providing transparent and interpretable decision-making processes. This paper explores the intersection of XAI and IIoT, highlighting the challenges of applying explainable models in real-time, data-intensive industrial contexts. We survey existing XAI techniques and evaluate their suitability for IIoT applications, such as predictive maintenance, quality assurance, and anomaly detection. Additionally, we discuss evaluation metrics, present case studies, and propose a framework for integrating XAI into IIoT pipelines. Our findings demonstrate the potential of XAI to enhance transparency, user trust, and operational safety in next-generation industrial systems.

Nandhini Ravi · 0 citations
Book Open access Jul 2026

Agents in the Wild: Where Research Meets Deployment

Through applied case studies in pharmaceutical discovery and financial systems, common design patterns that make agentic systems successful are analyzed, and practical mitigation strategies for failure modes are discussed, such as verification pipelines, fallback mechanisms, and human-in-the-loop supervision.

Grace Hui Yang, P. Venkit, Hooman Sedghamiz et al. · 0 citations
Open access 2025

Smarter AI Agents: Optimizing Tokens the Right Way

AI agents driven by large language models (LLMs) are radically changing industries through methods like automation, decision support, and intelligent interactions. As a result, the efficiency of these systems is as crucial as their capabilities. In fact, one of the most critical factors influencing AI's cost-effectiveness, speed, scalability, and user experience is token optimization, a factor often ignored in performance considerations. Some of the characteristics of the very modern AI workflows that can lead to a token explosion include deep prompting, several agents' interactions, retrieval of memories, and persistent context sharing. Such overuse of tokens has a double effect of continually increasing the expenditure and leading to unpleasant situations like lag, context overflow, deterioration of expected response, and wasting of resources. Alongside the contribution of AI agents to real-time applications, their critical nature is reminding us of the necessity to find ways of managing tokens intelligently to ensure a balance between performance and efficiency. This article presents a series of feasible and potent methods for optimizing token consumption in AI-driven systems at a minimum level without sacrificing the quality of outputs or the extent of contextual understanding. The methods proposed are prompt engineering antediluvian, context compression, memory selection, response generation, retrieval and adaptive token allocation that are area-specific and task-oriented and adapted to various workflows. Besides that, we analyze how intelligent token control can facilitate multi-agent collaboration operations while preventing unnecessary data exchange and reductions in processing redundancies. We maintain that enhancing token control has implications for a cleaner environment, greater scalability, and a more dependable and responsive infrastructure. The purpose of this paper is to present a practical, human-centric approach to the creation of 'smarter' AI agents, which are not only robust and precise but also resource-efficient and financially sustainable for large-scale deployment in the future.

Madhurima Kommuru · 0 citations
Open access Aug 2026

ASAREE: An Analytical Sandbox for Agentic AI Research, Engineering, and Experimentation

This work uses ASAREE to evaluate key design choices in a mutli-agent machine learning pipeline and highlights ASAREE as a robust framework for evaluating agentic system performance and resource efficiency.

Jay Moran, P. Freda, Attri Ghosh et al. · 0 citations
Open access Aug 2026

From Messy Data to Actionable Insights: Benchmarking Agentic AI Frameworks Across Real-World SHM Data Scenarios

The effective deployment of Artificial Intelligence (AI) in Structural Health Monitoring (SHM) is consistently hindered by the gap between idealized mathematical models and the field data, which is inherently heterogeneous, and asynchronous. While the industry has achieved high Technology Readiness Levels for isolated sensor hardware, it continues to suffer from low Integration Readiness Levels, forcing engineers into repetitive, manual data-wrangling tasks that preclude high-level diagnostic reasoning. This paper introduces a Reference Framework that proposes an architectural logic designed to navigate these integration challenges through Large Language Models (LLMs). We present a design hypothesis centered on a hybrid, multi-layered workflow that utilizes LLMs cognitive. As a result, this framework enables a cognitive layer to semantically interpret user intent and delegate complex, multi-step tasks to a library of deterministic, verified algorithms. To mitigate persistent integration risks and realize the potential of autonomous diagnostics, we propose that future SHM system designs should prioritize Pre-Processed Data Ingestion. By positioning the LLM as a bridge, this framework establishes a foundation for transitioning from manual data wrangling to conversational, actionable insights

Carlos Omar Rasgado moreno, Arjavi Salodkar, Aswin Haridas et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.