Skip to content
Preprint

RECTIFY: An Interactive Workbench for Post-Evaluation RAG Diagnosis, Repair, and Verification

Sep 2026 · 0 citations · 17 references
Computer Science

TL;DR

RECTIFY, an interactive Streamlit workbench that turns evaluated RAG cases into auditable repair workflows, is presented, showing that pre-filtering reduces unnecessary repair candidates and that slicelevel routing yields more targeted repair cards than broad family-level diagnosis.

Abstract

Retrieval-Augmented Generation (RAG) evaluators can identify failures such as weak retrieval, poor grounding, incomplete answers, and unsupported generation, but they rarely help developers decide what to repair next. We present RECTIFY, an interactive Streamlit workbench that turns evaluated RAG cases into auditable repair workflows. RECTIFY filters cases that do not require repair, routes remaining failures into actionable families and finegrained repair slices, and generates editable repair cards that developers can approve, reject, or verify through sandbox reruns. On a controlled RAG benchmark, RECTIFY surfaces interpretable failure profiles across BM25, dense, and hybrid retrieval: BM25 mainly triggers noisy-retrieval repairs, while dense and hybrid retrieval leave smaller sets of multi-part underretrieval and underused-evidence cases. Additional analyses show that pre-filtering reduces unnecessary repair candidates and that slicelevel routing yields more targeted repair cards than broad family-level diagnosis. RECTIFY is publicly available as an open-source Streamlit workbench 1 for helping developers turn evaluation results into inspectable repair decisions.

View source

Similar papers

Preprint Sep 2026

LENS: The Sum Is Worse Than the Parts for Set-Level Poisoning in Retrieval-Augmented Generation

Retrieval-augmented generation (RAG) aggregates evidence from multiple external documents, yet this joint integration creates an underexamined vulnerability: attack effects absent in individual documents can emerge through set-level composition. Existing coordinated attacks do not explicitly enforce that every proper s...

Kai-Sheng Fan, Yi-Shu Gao, Xun-Zhu Tang et al. · 0 citations
Conference Open access Sep 2026

IKnowFlow: Trustworthy RAG for Sensitive Domains

RAG systems ground LLM responses in external evidence, yet their trustworthiness remains underspecified. Retrieval, re-ranking, and generation are optimized independently with no unified account of whether the system is interpretable, grounded, and traceable. My thesis operationalizes trustworthy RAG through three curr...

Yash Saxena · 0 citations
Review Sep 2026

RCL: A Retrieval-Confidence Layer for Detecting Insufficient Context in Enterprise Retrieval-Augmented Code Generation

Retrieval-Augmented Generation (RAG) for code generation has been studied extensively on public repositories, where a model's parametric knowledge often compensates for imperfect retrieval. This breaks down in enterprise codebases, where private APIs, internal frameworks, and undocumented team conventions fall entirely...

Chandra Mohan Ravuri · 0 citations
#natural language process... Preprint Oct 2026

AGO AI Quality Gate: Evidence-First Release Decisions for Retrieval-Augmented Generation

Enterprises adopting retrieval-augmented generation (RAG) face a recurring operational decision: promote, revise, or block a system version. The evidence is incomplete and the metrics come from fallible LLM judges. We report on AGO AI Quality Gate (AGO), an evidence-first quality-gate framework deployed in industrial R...

Giulio Zeloni, Enrico Lo Conte, Salvatore Rionero et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.