Skip to content

RAILS: Retrieval-Augmented Incremental LLM Clustering at Scale

Sep 2026 · 0 citations · 27 references
Computer Science

TL;DR

RAILS is presented, a retrieval-augmented incremental LLM clusterer that turns clustering into a simple loop over a growing label pool and scales through document batching with bounded concurrency.

Abstract

Using a Large Language Model (LLM) as the clusterer at production scale is hard: prompts cannot hold the entire label space, and per-document serial processing does not deliver the throughput real workloads require. We present RAILS, a retrieval-augmented incremental LLM clusterer that turns clustering into a simple loop over a growing label pool and scales through document batching with bounded concurrency. On six public benchmarks RAILS exceeds the strongest prior LLM-clustering method on average, lifting accuracy from 51.2% to 59.3%, NMI from 67.2% to 74.8%, and ARI from 45.4% to 54.7%. We further report production-deployment evidence from a SaaS ticket-topic-discovery pipeline, where RAILS has replaced a traditional HDBSCAN stage with higher clustering quality, transparent prompt-driven control, and stateful incremental operation.

View source

Similar papers

Preprint Aug 2026

A Scalable Pipeline for LLM-Teacher Distillation Labeling: Work-Stealing Job Scheduling and Memory-Aware GPU Concurrency

Labeling large text corpora with LLM teachers has become a practical route to training data at scale. At millions of items, hand-labeling every batch is not feasible, and two questions dominate: what label quality a teacher buys per dollar, and how to keep a fleet of GPU workers busy under skewed, failure-prone workloa...

Ravi Satya Durga Prasad Yenugula · 0 citations
Book Open access Sep 2026

Batched in Back: Characterizing and Optimizing Offline LLM Inference in Production with ACDC

Serving offline large language model (LLM) inference workloads (e.g., log summarization and bulk translation) can consume up to 30% of GPUs in production. Despite this significant share, the characteristics of offline inference remain largely understudied. In this paper, we start by analyzing 1.5 million tasks comprisi...

Le-Ping Yang, Xue Li, Kun Qian et al. · 0 citations

Low-Latency Semantic Processing with Optimal Prompt-Level Batching

This paper measures and analytically derive an optimal size for in-prompt document batching to effectively amortize this overhead of LLM calls over public APIs, cutting end-to-end semantic operator latency by up to 14 × with no meaningful accuracy loss.

Jingyi Qu, Samuel Madden, Tianyu Li · 0 citations
Preprint Aug 2026

LLMVisor: A Real-Time Latency Attribution Model for Multi-Tenant LLM Serving

LLMVisor is presented, a roofline-guided latency attribution model that captures the memory-bound and compute-bound phases via a concise piecewise-linear form over features proportional to FLOPs and memory I/O traffic and runs efficiently at microsecond scale.

Shuowei Jin, Xue-Shen Liu, Jiaxin Shan et al. · 3 citations
Book Open access Sep 2026

LMTracer: Fine-Grained and Real-Time Performance Profiling for Production LLM Systems

LMTracer is presented, a fine-grained and real-time performance profiling framework for production LLM services that embed profiling logic into the execution through graph-embedded probing and streaming buffered profiling data to CPUs on demand to keep the execution of user kernels uninterrupted.

Wei Liu, Yong-Chao He, Bo-Han Zhao et al. · 0 citations
Preprint Sep 2026

TriCalRAG: A Three-Strategy, Retrieval-Augmented Benchmark for On-Premise LLM-Based Root Cause Analysis in AIOps

Operational logs create a need for private, resource-efficient incident analysis, but aggregate detection scores can conceal severe prediction bias. We present TriCalRAG, a reproducible benchmark for log-anomaly detection with generated root-cause and remediation outputs across BGL, HDFS, Thunderbird, and OpenStack. Th...

Rohit Patel, S. K. Mohanty, Jeenal Chaudhary · 1 citation

Related blog posts

MIT News · Artificial Intelligence Oct 7, 2026

Discovering the value of humanistic inquiry

Students in MIT’s Concourse program delve deeply into the human condition, debate challenging questions, and learn to develop judgment about issues that can’t be quantified.

Microsoft Research Blog Oct 7, 2026

Agent Lightning v1.0: A 3,500-Line Lightweight Agentic RL Framework for Training Agents with Real Harnesses

Training AI agents with reinforcement learning can be challenging because their tools, context, and decision-making are managed by complex frameworks. Agent Lightning connects existing agents to RL training, making it easier to improve them without rebuilding them. The post Agent Lightning v1.0: A 3,500-Line Lightweight Agentic RL Framework for Training Agents with Real Harnesses appeared first on Microsoft Research.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.