Aug 2026· Multimedia tools and applications· Vol 85· 0 citations· 42 references
TL;DR
This paper proposes RAG-Test, a developer-centric and unified framework that automates test case generation, attribution validation, and productivity estimation and introduces a theoretical productivity gains model, estimating chatbot efficiency improvements over search engines.
RagTester is presented, an automated end-to-end testing approach for RAG systems that can effectively expose failures caused by the interaction between retrieval and generation components and support the assessment of RAG configurations before deployment.
Ange Maiztegi, J. Ayerdi, Miren Illarramendi et al.· 0 citations
Findings indicate that clear prompt instructions can improve the reliability of RAG-based academic chatbot responses for academic information services and show the strongest improvement in faithfulness and context recall.
Adi Surya Artayasa, Aniek Suryanti Kusuma, Putu Sugiartawan· IJOEM: Indonesian Journal of...· 0 citations
XREPOTEST is introduced, a multilingual repository-level benchmark for unit test generation spanning five underexplored languages: Rust, Go, Julia, PHP, and Ruby, and Invocation Rate is proposed to assess whether generated tests meaningfully exercise the intended functionality.
L. Dung, Dong Cao Van, Nam Le Hai et al.· 0 citations
A reproducible, human-validated evaluation framework applied to 13 strategies—four architectural families crossed with four reasoning variants crossed with four reasoning variants—across three SLMs spanning 3B–14B parameters, plus targeted ablations.
Balaji Venktesh, Amsaprabhaa M, G. Sundaram· International Conference on...· 0 citations
We introduce SciRet, a compute-aware empirical study of retrieval-augmented generation for scientific question answering over CORD-19. Rather than proposing a new model, we evaluate a fixed scientific RAG pipeline across three corpus scales: 1,034 chunks (1K papers), 5,160 chunks (5K papers), and 15,480 chunks (15K papers). The pipeline combines sentence-window chunking, BM25, BGE-M3 dense retrieval, reciprocal rank fusion, optional cross-encoder reranking, and grounded answer generation. Across these settings, hybrid retrieval is more robust than either sparse-only or dense-only retrieval in our setting, reaching Recall@10 of 1.000 at 1K and 15K. In contrast, an MS MARCO-trained cross-encoder reranker reduces precision on the scientific corpus, suggesting that domain mismatch can outweigh the benefits of stronger query-passage interaction. Generation faithfulness measured with RAGAS increases with corpus scale in our setup. Retrieval evaluation uses pseudo-relevance labels derived from the hybrid system, so we treat the results as controlled comparative evidence rather than a benchmark claim. We release code, indexes, and evaluation outputs to support replication and follow-up studies.
Kaysarul Anas Apurba, Mahade Hasan, Rofiqul Alam Shehab et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.