RECTIFY, an interactive Streamlit workbench that turns evaluated RAG cases into auditable repair workflows, is presented, showing that pre-filtering reduces unnecessary repair candidates and that slicelevel routing yields more targeted repair cards than broad family-level diagnosis.
Abstract
Retrieval-Augmented Generation (RAG) evaluators can identify failures such as weak retrieval, poor grounding, incomplete answers, and unsupported generation, but they rarely help developers decide what to repair next. We present RECTIFY, an interactive Streamlit workbench that turns evaluated RAG cases into auditable repair workflows. RECTIFY filters cases that do not require repair, routes remaining failures into actionable families and finegrained repair slices, and generates editable repair cards that developers can approve, reject, or verify through sandbox reruns. On a controlled RAG benchmark, RECTIFY surfaces interpretable failure profiles across BM25, dense, and hybrid retrieval: BM25 mainly triggers noisy-retrieval repairs, while dense and hybrid retrieval leave smaller sets of multi-part underretrieval and underused-evidence cases. Additional analyses show that pre-filtering reduces unnecessary repair candidates and that slicelevel routing yields more targeted repair cards than broad family-level diagnosis. RECTIFY is publicly available as an open-source Streamlit workbench 1 for helping developers turn evaluation results into inspectable repair decisions.
Retrieval-augmented generation (RAG) aggregates evidence from multiple external documents, yet this joint integration creates an underexamined vulnerability: attack effects absent in individual documents can emerge through set-level composition. Existing coordinated attacks do not explicitly enforce that every proper s...
Kai-Sheng Fan, Yi-Shu Gao, Xun-Zhu Tang et al.· 0 citations
It is proved that, under an honest-majority assumption and a representation-level separation condition, RAGSentinel exactly recovers a poison-free majority-sized context.
Yueyang Quan, Anjun Gao, Yu Xia et al.· 0 citations
RAG systems ground LLM responses in external evidence, yet their trustworthiness remains underspecified. Retrieval, re-ranking, and generation are optimized independently with no unified account of whether the system is interpretable, grounded, and traceable. My thesis operationalizes trustworthy RAG through three curr...
Yash Saxena· Proceedings of the Thirty-Fi...· 0 citations
Retrieval-Augmented Generation (RAG) for code generation has been studied extensively on public repositories, where a model's parametric knowledge often compensates for imperfect retrieval. This breaks down in enterprise codebases, where private APIs, internal frameworks, and undocumented team conventions fall entirely...
Enterprises adopting retrieval-augmented generation (RAG) face a recurring operational decision: promote, revise, or block a system version. The evidence is incomplete and the metrics come from fallible LLM judges. We report on AGO AI Quality Gate (AGO), an evidence-first quality-gate framework deployed in industrial R...
Giulio Zeloni, Enrico Lo Conte, Salvatore Rionero et al.· 0 citations
This survey presents a unified and pipeline-aware overview of RAG robustness, formalize threat models over the corpus, retriever, and generator, and organize attacks into three main objectives: accuracy, privacy, and fairness.
Minh Tran, Cuong Dang, Tuc Nguyen et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.