A Reliability-Aware Retrieval-Augmented Generation Architecture for Open Language Models in Higher Education Decision Support
Abstract
Open language models are increasingly considered for institutional decision-support tasks in higher education, including policy interpretation, academic advising, administrative summarization, and quality-assurance workflows. However, their reliable deployment requires more than model availability: it depends on cloud-native orchestration, retrieval quality, evidence grounding, refusal behavior, monitoring, and governance controls. Following a design-science research approach, this paper presents an architectural artifact for deploying open language models in higher education decision support. The artifact operationalizes institutional reliability as a multidimensional construct composed of contextual accuracy, answer faithfulness, retrieval quality, refusal adequacy, latency compliance, auditability, and human-review compatibility, and aggregates these into an institutional reliability index. It proposes a reliability-aware retrieval-augmented generation pipeline that integrates governed document ingestion, embedding generation, hybrid retrieval, reranking, evidence-aware generation, confidence-based refusal, human review, audit logging, and post-deployment monitoring. To support reproducibility, the paper compares four deployment configurations and provides an illustrative worked example of the reliability index. The contribution is a conceptual yet technically grounded deployment artifact that connects cloud computing, data science, and higher education governance; the architecture has not yet been empirically validated, and a protocol for future institutional pilots is specified.