Skip to content
Review

A Semi-Automated System for Generating Dialogue-Based TTS Lessons Using Large Language Models: An Exploratory Study of Educational Potential

Jul 2026 · 1 citation · ⚡ 1 influential · 24 references
Computer Science

TL;DR

A theoretical and empirical basis is provided for the educational acceptability of TTS audio and for TTS lesson-format design and a trade-off between dialogue's benefits and higher extraneous cognitive load is revealed.

Abstract

This study proposes a semi-automated system for generating dialogue-based lessons using Large Language Models (LLMs) and Text-to-Speech (TTS) technology, and exploratorily examines its educational potential via a practical quasi-experiment. The system augments rather than replaces educators through a three-stage human-in-the-loop workflow (LLM-based slide/narration generation, educator review, automated audiovisual integration), and introduces a novel method for generating Expert-Novice dialogue narration based on cognitive apprenticeship theory. In a study of 245 first-year high school students who sequentially experienced three lesson formats (instructor voice, single-speaker TTS, dialogue TTS; content differed across sessions, limiting format/content separation), we conducted within-subject (Friedman test, N<=183) and repeated cross-sectional (Mann-Whitney U, N=229/206) analyses. TTS audio did not substantially degrade the learning experience versus instructor voice, supported by TOST equivalence testing. Dialogue TTS was significantly superior to single TTS in comprehension (p=.006, q=.025) and cognitive engagement (p=.019, q=.048); enjoyment was non-significant after FDR correction (q=.081) but reached significance after controlling for prior knowledge (proportional-odds model, OR=1.65, q=.025), and these advantages were not attributable to prior-knowledge imbalance. Conversely, single TTS was superior in audio naturalness (p<.001, q<.001, r=-.238), revealing a trade-off between dialogue's benefits and higher extraneous cognitive load. Dialogue format was preferred by 66.9% of learners as most enjoyable (p<.001). These results reflect a fixed-order design; replication is needed before generalizing them as effects of lesson format. This study provides a theoretical and empirical basis for the educational acceptability of TTS audio and for TTS lesson-format design.

View source

Similar papers

#small language model Preprint Aug 2026

Interaction Effects Between Learner Characteristics and Dialogue Format in TTS Dialogue-Based Lessons

The results suggest that dialogue format should be selected according to learner characteristics in TTS dialogue-based lessons, with a significant interaction between learner characteristics and dialogue format for ARCS-based motivation.

Fumie Watanabe, Tota Suko, T. Ishida et al. · 0 citations
Open access Jul 2026

Evaluating the Validity of a Spoken Dialog System-Mediated Role-Play Test for Assessing Second Language Oral Communication

Generative AI offers new opportunities for interactive oral communication assessment for specific purposes by building on past research investigating test takers’ oral interaction with spoken dialog systems (SDSs). Performing as automated conversational agents, SDSs can address logistical challenges in classroom-based oral assessments where interactive tasks are difficult to implement. This study investigated the validity of interpretations of a prototype SDS-mediated Tourism English Speaking Test (SDS-TEST) developed using evidence-centered design. Guided by the interpretation/use argument framework, this study focused on supporting the generalization and explanation inferences. Thirty Turkish English as a Foreign Language (EFL) students in a Tourism and Hotel Management program completed three SDS tasks, each rated by four trained raters. The generalization inference was supported using a univariate G-study with a fully crossed, two-facet design on 360 composite scores, while a D(ecision)-study informed optimal task and rater configurations for future designs. The explanation inference was supported through correlational analyses with related assessments and qualitative insights from participants’ strategy use gathered via stimulated recalls. Findings support the inferences in the validity argument for the SDS-TEST, highlighting the potential of SDSs in English for Specific Purposes oral assessment and offering a model for exploring validity in the next generation of oral communication assessment.

Yasin Karatay, C. Chapelle · 0 citations
Conference Jul 2026

An On-Premise Multilingual Academic Chatbot using Retrieval-Augmented Generation and Context-Aware Memory for University Assistance

Universities now use Large Language Models (LLMs) to transform their processes for managing student information. The paper introduces an upgraded chatbot system for Narasaraopeta Engineering College (NEC) which extends previous on-premise LLM chatbot research by providing four new functions. The system uses (1) Retrieval-Augmented Generation (RAG) to create citation-based responses through LlamaIndex and ChromaDB, (2) Context Memory which maintains conversation flow during multiple dialogue exchanges, (3) Voice Input through OpenAI Whisper Speech-to-Text (STT) technology, and (4) Multilingual Support which covers English and these seven languages: Hindi, Telugu, Tamil, Kannada, and Malayalam through IndicNLP. The system tested 60 benchmark questions across four academic categories which included regulations and examination policies and fee structures and multilingual queries and achieved 96.7% overall accuracy with sub-second text response times and 1.0–1.4 second voice response times. The system operates entirely on-premise through Docker which safeguards institutional data privacy while eliminating the need for recurring cloud API expenses. The upcoming development will create Emotion-Aware AI, FAQ Auto-Learning, Student Portal Integration, and a Mobile Application.

M. Yaswanth, Kopparapu Sai Amar Durgesh, Mogili Harsha Vardhan et al. · 0 citations
Book Open access Aug 2026

PLAI: A Pilot Study of Profile-Based Explanation for AI-Supported Learning

Large language models (LLMs) are increasingly used as on-demand conversational learning assistants, but they typically do not adapt explanations to a student’s background unless explicitly prompted. We present the Personalized Learning Assistant Interface (PLAI), a web-based prototype that generates explanations from lecture slides, audio transcripts, and a structured student profile through a chat-based interface. We evaluated PLAI in a controlled pilot study with 24 STEM students, comparing profile-based personalization with a baseline condition using the same slide and transcript context. Immediate learning was assessed with a five-item knowledge test, while subjective experience was measured using the User Experience Questionnaire (UEQ) and open-ended feedback. We did not observe clear differences in knowledge-test outcomes, but participants in the personalized condition reported significantly higher UEQ Stimulation. These results suggest that profile-based multimodal prompts may mainly support motivational engagement rather than immediate test performance, while larger and longer-term studies are needed to assess learning effects.

Furkan Ali Yurdakul, Yiman Wu, Maria Torres Vega et al. · 0 citations
Review Open access 2026

Beyond the AI Conversation Partner: A Critical Integrative Review of Generative AI-Supported L2 Speaking and the SPEAK Implementation Framework

Generative artificial intelligence (GenAI) has rapidly expanded from text-based assistance to voice-enabled dialogue, automated oral feedback, and multimodal speaking assessment. These developments appear to address a persistent problem in second-language (L2) education: learners need frequent, low-risk opportunities to speak, but teachers cannot always provide individual interaction and feedback at scale. However, the availability of an apparently fluent conversation partner does not establish that durable speaking development, fair assessment, or transfer to human communication will follow. This critical integrative review examines three questions: what learning and affective outcomes are associated with GenAI-supported L2 speaking practice; what limitations weaken its pedagogical value; and what implementation conditions support responsible use. Targeted searches of Google Scholar, publisher platforms, open research repositories, and citation chains produced an analytical corpus of 11 core publications, supported by 14 theoretical, methodological, assessment, feedback-literacy, and ethics sources published or retained for interpretation through July 2026. Evidence was appraised for contextual clarity, task alignment, outcome validity, feedback transparency, and the strength of claims. The synthesis indicates that GenAI can increase practice volume, support rehearsal, provide scenario-based interaction, and reduce fear of immediate human judgement. Adaptive prompts and rapid feedback may also support vocabulary retrieval, discourse organisation, pronunciation awareness, and willingness to communicate. Yet the evidence remains dominated by small samples, short interventions, self-report measures, prototype studies, and technical benchmarks. Recurring risks include inaccurate or overconfident feedback, accent and speech-recognition bias, unnatural interaction, dependency, privacy concerns, unequal access, and weak transfer evidence. The review proposes the SPEAK framework: Structured tasks, Progressive intelligibility-focused feedback, Equity and ethics, Agency and anxiety-sensitive practice, and Knowledgeable teacher oversight. GenAI should therefore be used as a structured rehearsal and feedback resource within teacher-designed speaking pedagogy, not as an autonomous replacement for human interaction or professional judgement.

D. Dasanayake · 0 citations
Preprint Aug 2026

TACT: Taxonomy-Aligned Post-Training for Pedagogically Adaptive English Tutoring

Large language models (LLMs) are increasingly used to provide conversational practice for English-as-a-second-language (ESL) learners. Effective ESL tutoring, however, requires more than fluent response generation: a tutor must select an appropriate pedagogical action based on learner behavior and dialogue context. Human-tutoring research offers principles for adaptive support, but they are often task-specific and remain insufficiently integrated into LLM-based ESL tutor training and evaluation. We present TACT (Taxonomy-Aligned Conversational Tutor), a human-grounded framework for post-training and evaluating pedagogically adaptive ESL tutors. Drawing on established literature, we develop two complementary taxonomies: the Tutor-Strategy Taxonomy with 13 tutor response strategies and the Student-Move Taxonomy characterizing learner behavior by move type and status. Using these taxonomies, we construct TACTCorpus, which enriches 260 authentic teacher-student conversations with 32,379 annotations and quality-controlled augmented training data. We then post-train Qwen3.5-4B through supervised fine-tuning followed by taxonomy-aligned Group Relative Policy Optimization, producing TACTutor and optimizing it for scaffolding quality rather than reference imitation alone. On TACTBench, a strategy-balanced diagnostic benchmark comprising 78 authentic tutoring contexts, TACTutor improves over its backbone by 20.30% and outperforms all evaluated proprietary baselines under the same protocol, while maintaining backbone performance on established external educational benchmarks; in a blinded study with 50 learners, it also receives the highest overall mean rating among the evaluated tutors. We release the data, benchmark, and model weights, providing an open foundation for developing pedagogically adaptive ESL tutors.

Dongjie Yang, Siya Lin, Leixian Shen et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.