RAG-Powered Clinical Conversational Agent for Guideline Retrieval: A Prototype System Comparing Standard and HyDE-Enhanced Retrieval
Abstract
Healthcare professionals spend significant time manually retrieving information from clinical guidelines during time-sensitive scenarios. This study presents a prototype conversational AI agent that leverages Retrieval-Augmented Generation (RAG) to enable rapid, natural-language access to obstetric clinical guidelines. The system was built using n8n as the workflow orchestration platform, Supabase as the vector database, GPT-4o as the generative model, and OpenAI text-embedding-3-small for document embeddings. A corpus of 18 obstetric clinical guidelines was ingested, yielding approximately 3721 text chunks. We compared two retrieval strategies: standard RAG and RAG enhanced with Hypothetical Document Embedding (HyDE), where the latter generates hypothetical answers to guide the retrieval process. Both configurations were evaluated on 39 clinical questions with ground truth answers provided by six obstetric specialists, using RAGAS metrics: faithfulness, answer relevancy, context precision, and context recall, computed via an LLM-as-a-judge pipeline. Results show that HyDE-enhanced RAG achieved marginal improvements in faithfulness (0.69 vs. 0.64) and context precision (0.63 vs. 0.58), while standard RAG showed higher answer relevancy (0.70 vs. 0.66). Both approaches obtained identical context recall (0.50). Overall, the system demonstrates the feasibility of RAG-based agents for obstetric guideline retrieval with information traceability to source documents, though further validation with clinical end-users is needed before deployment.