This paper proposes ThemePath-RAG, a retrieval framework that retrieves curated thematic paths as high-recall semantic routes, expands candidate canonical evidence, and applies query-aware scoring and global pruning before generation.
Abstract
Retrieval-Augmented Generation (RAG) grounds large language models in external evidence, but many RAG systems represent knowledge either as flat text chunks or as automatically constructed indexing graphs. This assumption is incomplete for curated thematic corpora, including religious scriptures, legal codes, clinical guidelines, educational taxonomies, policy documents, and library classification systems, where domain experts have already organized knowledge into thematic paths and citeable canonical units. This paper investigates how RAG can exploit such expert-authored structures while pruning evidence to a compact and query-specific set. We conduct a critical survey supported by a bibliometric analysis of 2815 Scopus-indexed RAG-related records exported on 26 May 2026, of which 2809 records were retained after duplicate removal. The bibliometric results indicate rapid growth in RAG research but limited explicit consolidation around curated thematic paths, canonical evidence units, or thematic path-guided evidence pruning. We therefore propose ThemePath-RAG, a retrieval framework that retrieves curated thematic paths as high-recall semantic routes, expands candidate canonical evidence, and applies query-aware scoring and global pruning before generation. To assess operational feasibility, we implement ThemePath-RAG for Qur’anic question answering and compare it with a Vector RAG baseline on 150 paired questions using RAGAS context relevance with gpt-4o-mini as the LLM evaluator. Both methods return approximately three final ayat per question. Vector RAG achieves higher mean context relevance than ThemePath-RAG (0.920 versus 0.798; p<0.001). Thus, the proof of concept establishes the feasibility of thematic-path-guided retrieval and identifies evidence-selection challenges, rather than demonstrating superiority over conventional vector retrieval. The paper clarifies the framework’s relationship to GraphRAG, LightRAG, HippoRAG, PathRAG, ontology-based RAG, and AI-augmented bibliometric systems, and outlines a language-matched, multi-baseline evaluation agenda for future cross-domain validation.
Although multi-turn inference remains more expensive than single-call retrieval, VecTree-RAG provides a structure-aware and traceable architecture for scientific literature question answering.
The evidence indicates that no single RAG or vector-database configuration dominates across retrieval quality, faithfulness, latency, throughput, storage, cost, and scalability, and the review positions RAG–vector database integration as a joint retrieval-and-systems optimization problem rather than a database-selection problem alone.
Muhammad Fuad Bin Abdullah, Safwan Abd Razak, Noorrezam Yusop et al.· International journal of res...· 0 citations
This work proposes a pioneering hierarchical structure-retrieved generation framework (HS-RAG) that reconceptualizes the generation task as a systematic retrieval-alignment-fusion process from template to text, marking a fundamental paradigm shift from spontaneous generation to grounded structural anchoring.
Yongpan Wang, Yu Tan, Mingli Song et al.· IEEE Access· 0 citations
W-RAG is proposed, a source-aware retrieval framework that performs ontology-guided retrieval, local ranking within each knowledge base, and source-level weighting to regulate evidence composition to improve document coverage and generation quality.
Hridya Dhulipala, Rajesh Ombase, Michael Wang et al.· 0 citations
Objective. The objective of this study was to provide a comprehensive bibliometric review of research at the intersection of retrieval-augmented generation (RAG) and knowledge graphs (KGs), to map the intellectual structure of the field, and to quantify scholarly activity on system architectures, evaluation venues, and emerging application domains from 2021 to 2026.
Design/Methodology/Approach. A bibliometric review was conducted in accordance with the PRISMA 2020 reporting framework. A combined Scopus and Web of Science (WoS) search yielded 1,604 records (Scopus, n = 1,098; WoS, n = 506). After the removal of 444 cross-database duplicates using the KKU-BiblioMerge toolkit and the screening of 1,604 unique records against four exclusion criteria, 313 records were excluded, leaving 1,291 publications for analysis. A variety of bibliometric techniques were employed in the analysis, including descriptive bibliometrics, co-authorship analysis, country collaboration mapping, keyword co-occurrence, thematic mapping, co-citation analysis, and latent Dirichlet allocation topic modeling (k = 8, u_mass coherence = −1.46). The analyses were conducted in Python, utilizing the libraries pandas, NetworkX, and gensim.
Results/Discussion. The findings indicated a rapid expansion of the field, with annual publication output increasing from 1 paper in 2021 to 700 papers in 2025 (average annual growth rate (AGR) of 235.49%). When author affiliations were aggregated at the unique-publication level and the two WoS variants “China” and “Peoples R China” were merged, Mainland China (n = 485 publications) and the United States (n = 139) emerged as the leading contributors. The most-cited publication was “Knowledge Graph Prompting for Multi-Document Question Answering,” with 110 citations. A total of eight major thematic clusters were identified, with “Retrieval-Augmented Generation Systems” and “Embedding & Semantic Representation” demonstrating the highest average citation impact. The findings suggested a growing scholarly focus on GraphRAG architectures and KG-enhanced large language model pipelines to enhance factual accuracy and knowledge grounding.
Conclusions. Research on the integration of RAG and KGs is rapidly maturing and consolidating around hybrid architectures that combine structured knowledge retrieval with neural text generation. However, significant challenges persist, particularly regarding standardized evaluation benchmarks, cross-lingual RAG–KG systems, and scalable metadata generation frameworks.
Originality/Value. This study offers the inaugural comprehensive bibliometric review of the RAG–KG research landscape. It provides a comprehensive intellectual framework for the field, identifies emerging research directions, and offers practical insights for the development of intelligent metadata-generation assistants and knowledge-enhanced AI systems.
It is suggested that, in dense-urban POI settings where coordinates are reliable, the marginal benefit of explicit graph edges shrinks for coordinate-computable relationships, while structured spatial processing complements vector retrieval.
Noboru Otsuka· The International Archives o...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.