This work systematically evaluates edge-oriented ASR-LLM pipelines for individuals with language impairments using comparison studies and ablation experiments across aphasia, child language impairment, and dementia datasets to identify transcript errors, repetition, noise, and input length as factors affecting system performance and deployment feasibility.
Abstract
Large language models (LLMs) have attracted much attention for healthcare applications, demonstrating strong potential in automating conversational interactions. However, cloud-hosted systems raise privacy concerns and text-based interfaces can be difficult for children, older adults, and individuals with language impairments. Edge-deployed, voice-enabled LLM systems may address these barriers by keeping data local and using automatic speech recognition (ASR) for speech input. However, ASR transcripts often contain disfluencies, fillers, grammatical errors, and recognition noise, which may reduce downstream LLM performance for speakers with atypical language patterns. Here, we systematically evaluate edge-oriented ASR-LLM pipelines for individuals with language impairments using comparison studies and ablation experiments across aphasia, child language impairment, and dementia datasets. We identify transcript errors, repetition, noise, and input length as factors affecting system performance and deployment feasibility. These findings highlight key challenges for building more reliable and accessible speech-enabled AI systems for healthcare applications.
This work analyzed brief story retellings from 86 patients with left-hemisphere stroke and derived discrete linguistic features and embeddings with Large Language Models, providing proof of concept for a fast, largely automated discourse screener of acute LI.
Automatic language proficiency assessment in the context of multilingual interview-based settings remains underexplored. In this work, we develop Whisper-based speaker-role and language diarization systems to automatically extract respondent speech and characterize language usage in multilingual interviews with older a...
Anfeng Xu, Tian-Tian Feng, Kevin Huang et al.· 0 citations
Aphasia is a neurological condition that affects an individual's ability to speak and understand language. Accurate analysis and evaluation of patient speech are essential for effective rehabilitation. In this study, an artificial intelligencebased speech therapy system is proposed to support aphasia rehabilitation. Sp...
S. Devasurithi, M. Madhumitha, G. Nandhini et al.· International Conference on...· 0 citations
Recent advances in Speech Language Models (SpeechLMs), which integrate large language models with speech foundation models, have enabled unified sequence modeling of speech processing tasks. However, many SpeechLM-based approaches to speaker diarization (SD) are tightly coupled with automatic speech recognition (ASR) a...
CASA, a simpler architecture combining Whisper-medium and Qwen3.5-2B that achieves state-of-the-art performance while providing a more interpretable separation between speech delivery and content, is proposed.
Nhan Phan, Ilona Lähteenmäki, Anna von Zansen et al.· 0 citations
Child speech differs from adult speech in acoustics, prosody, and linguistic structures. Speech disfluencies (such as repetitions) further challenge automatic understanding. While Audio Language Models (ALMs) show strong semantic reasoning from speech audio, their ability to reason about disfluent child speech in mixed...
C. Okocha, Christan Grant, Zoey Liu· 1 citation
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.