Skip to content
Open access

EduAssist: Evaluating a Locally Deployed Large Language Model for Educational Document Summarization

Aug 2026 · International Journal of Engineering Research and Science & Technology · Vol 22, pp. 1315-1322 · 0 citations

TL;DR

E EduAssist, an experimental framework for evaluating a locally deployed Qwen3:1.7B model for educational document summarization, provides exploratory empirical evidence for the feasibility of local educational document summarization using the evaluated Qwen3:1.7B configuration while highlighting the need for larger datasets, model comparisons, independent reference summaries, and human evaluation.

Abstract

The increasing use of large language models (LLMs) for educational text processing has created opportunities for automatic summarization of lengthy learning materials. However, many LLM-based applications rely on cloud-hosted services, while the performance and computational behavior of locally deployed language models for educational document summarization remain comparatively underexplored. This study presents EduAssist, an experimental framework for evaluating a locally deployed Qwen3:1.7B model for educational document summarization. The model was executed through the Ollama runtime using a chunk-based summarization and consolidation pipeline. Experiments were conducted on 12 English-language educational samples covering topics in data mining, machine learning, artificial intelligence, and knowledge-based systems. Evaluation combined text-reduction and efficiency measures with lexical, semantic, and model-assisted assessment using compression ratio, estimated reading-time reduction, processing time, source-reference ROUGE, BERTScore, and LLM-as-a-Judge. The generated summaries achieved a mean compression ratio of 50.83%, reducing the mean source length from 1,074.17 to 478.42 words and yielding an estimated mean readingtime saving of 2.98 min. Mean ROUGE-1, ROUGE-2, and ROUGE-L scores were 0.4917, 0.2371, and 0.2892, respectively, while the mean BERTScore F1 was 0.8264 and the mean LLM-as-a-Judge score was 7.96/10. Summary generation required an average of 194.33 s per sample. Exploratory analysis further showed that greater compression was associated with lower source-reference lexical retention, whereas source-summary BERTScore F1 values remained comparatively stable across the evaluated samples. Overall, the findings provide exploratory empirical evidence for the feasibility of local educational document summarization using the evaluated Qwen3:1.7B configuration, while highlighting the need for larger datasets, model comparisons, independent reference summaries, and human evaluation.

Read PDF

Similar papers

Review Open access Jul 2026

Intelligent One-Gate System Based on Natural Language Processing for Enhancing Academic Information Services

An integrated transformer-based one-gate architecture that combines document routing, academic document summarization, and conversational assistance in a single service platform, offering practical value for reducing fragmented academic information services and methodological value as a reference model for higher education NLP implementation is developed.

Jaka Purnama, Yayuk Ike Meilani · 0 citations
Open access Jul 2026

Implementation of Transfer Learning for Automatic Summarization in Research Article Synthesis

The increasing number of scientific publications has made literature screening more time-consuming, particularly for researchers who need to identify the main contribution of an article before reading the full text. This study develops an extractive summarization model for Indonesian scientific articles using a BERT-based transfer learning approach. The proposed method represents each sentence with contextual embeddings and selects relevant sentences based on their similarity to the document representation, while applying a redundancy threshold to reduce redundancy. A curated corpus of Indonesian research articles was used for model development and evaluation. The generated summaries were evaluated using ROUGE-1, ROUGE-2, ROUGE-L, and ROUGE-Lsum, with the article abstract used as the reference summary. The experimental results show that the proposed model achieved a ROUGE-1 score of approximately 33%, indicating it retained important keywords and central information from the source documents. However, the lower ROUGE-2 score suggests that the model still has limitations in preserving phrase-level continuity and sentence coherence. Qualitative analysis also shows that the model can capture the main ideas of scientific articles, although some methodological details and contextual information are occasionally omitted. These findings indicate that BERT-based extractive summarization can support preliminary literature screening, but further improvement is needed through stronger baseline comparison, human evaluation, and redundancy-aware optimization.

Made Hanindia Prami Swari, Puji Lestari Tarigan, Gusti Eka Yuliastuti et al. · 0 citations
2026

Assessing Transformer Models for Abstractive Summarization of Scientific Articles

The results show that BART achieves the best performance with an ROUGE-2 F1-score of 0.40664, while T5 demonstrates superior grammatical acceptability, achieving 93.36%, but BART achieves a very near performance to T5.

Emad Nabil · 0 citations
Open access Jul 2026

AI-Based Document Analysis and Question Answering System

This study provides an AI- Based document analyzer with a question-answer system that makes use of Natural Language Processing approaches that is affordable, scalable, and suitable for business, education, and research.

Radhika Sharma, Devraj Gautam · 0 citations
Open access 2026

Leveraging large language models for scalable analysis of the end-of-the-course student feedback

Analyzing open-ended student feedback in course evaluations is a laborintensive task due to the unstructured and complex nature of natural language. While Large Language Models (LLMs) offer significant potential for automation, a welldefined methodology for their application in analyzing student feedback remains underdeveloped. This paper addresses this gap by proposing an LLM-based feedback analytics pipeline designed to transform students’ open-ended feedback into structured, actionable insights. The pipeline consists of three sequential stages: (i) segmenting student feedback into semantic units and assigning polarity (sentiment) to those units; (ii) topical classification of semantic units, and (iii) summarization of units within each topical category. By systematizing these processes, the proposed method enables educators and course managers to efficiently derive meaningful patterns from vast datasets of student opinions. We evaluated the proposed method using a comprehensive dataset from several editions of a U.S. university course, yielding encouraging results of the method’s effectiveness. This research provides a scalable, generic methodology for (semi-)automated feedback analysis, ultimately supporting data-informed improvements in teaching and course management.

Unknown authors · 0 citations
Open access Aug 2026

An Optimization Framework for Retrieval Augmented Generation in Indonesian Educational Question Answering

A RAG optimization framework for Indonesian-language educational question answering using a Human-Computer Interaction learning corpus as a case study is developed and provides a procedure for selecting retrieval and generation settings for a given corpus.

I. K. R. Arthana, N. Gunantara, Made Sudarma et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.