Open language models are increasingly considered for institutional decision-support tasks in higher education, including policy interpretation, academic advising, administrative summarization, and quality-assurance workflows. However, their reliable deployment requires more than model availability: it depends on cloud-native orchestration, retrieval quality, evidence grounding, refusal behavior, monitoring, and governance controls. Following a design-science research approach, this paper presents an architectural artifact for deploying open language models in higher education decision support. The artifact operationalizes institutional reliability as a multidimensional construct composed of contextual accuracy, answer faithfulness, retrieval quality, refusal adequacy, latency compliance, auditability, and human-review compatibility, and aggregates these into an institutional reliability index. It proposes a reliability-aware retrieval-augmented generation pipeline that integrates governed document ingestion, embedding generation, hybrid retrieval, reranking, evidence-aware generation, confidence-based refusal, human review, audit logging, and post-deployment monitoring. To support reproducibility, the paper compares four deployment configurations and provides an illustrative worked example of the reliability index. The contribution is a conceptual yet technically grounded deployment artifact that connects cloud computing, data science, and higher education governance; the architecture has not yet been empirically validated, and a protocol for future institutional pilots is specified.
I. García-López, Nicia Guillén-Yparrea· Computers· 0 citations
This study examines whether open large language models (OLLMs) can be optimized for equitable and efficient use in higher education under resource-constrained conditions. A quantitative experimental design was implemented to evaluate five OLLMs —Falcon, Bloom, GPT-NeoX, T5, and Flan-T5—under four conditions: baseline unoptimized inference (C0), pruning only (C1), retrieval-augmented generation (RAG) only (C2), and pruning combined with RAG (C3). Each condition was tested using 50 educational queries per model across four repetitions. To justify the pruning configuration, an ablation study compared sparsity levels of 10%, 20%, 30%, and 40%, identifying 20% as the best trade-off between efficiency and response quality. The results show that pruning reduced response time by 10.7%, lowered RAM/VRAM usage by 18.9%, and increased throughput by 33.2%. Retrieval augmentation improved educational response accuracy by 1.3 percentage points, with the strongest gains observed for factual queries. Although pruning-only achieved the best efficiency results and RAG-only produced the highest accuracy, the combined condition provided the most balanced profile for realistic educational deployment. These findings suggest that moderate pruning and retrieval augmentation can jointly support lighter, faster, and more contextually grounded language-model deployment in higher education, particularly in institutions with limited computational infrastructure. The study contributes empirical evidence of technical and pedagogical feasibility under simulated deployment conditions.
I. García-López, J. Molina-Espinosa, M-S. Ramirez-Montoya· IEEE Revista Iberoamericana...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.