A Comparative Study on Knowledge Graph-based RAG Models for Improving Multi-hop Reasoning in Administrative Regulations: Focusing on University Regulation QA Systems
Abstract
This study constructs and comparatively analyzes Retrieval-Augmented Generation (RAG) and GraphRAG (RAPTOR RAG, Community-GraphRAG, HippoRAG2) systems based on university regulations to enhance administrative efficiency and knowledge management. Utilizing the Llama-3.1-8B language model and the text-embedding-ada-002 embedding model, we evaluated the systems across single-hop and multi-hop reasoning queries. To comprehensively assess retrieval and generation performance, diverse evaluation metrics were applied, including Recall@5, F1-score, Faithfulness, Answer Relevance, and Latency. The quantitative results demonstrated that while RAG offered efficient latency (avg. 2.1s) for single-hop factual retrieval, HippoRAG2 showed superior performance in multi-hop reasoning environments, achieving a Faithfulness of 0.942 and Answer Relevance of 0.892 without the inclusion of total accuracy metrics. Based on these empirical comparisons, this study proposes a hybrid RAG strategy as a practical guideline for advancing customized AI knowledge search systems in universities.