Skip to content
Review Open access

Voice, Speech, and Large Language Models in Neurology: From Acoustic Biomarkers to Conversational AI

Jul 2026 · De Computis · Vol 14, pp. 160 · 0 citations · 131 references
Computer Science

TL;DR

The field is moving toward integrated speech-language assessment, but the gap between technical capability and clinical utility remains wide, and closing it requires diverse multilingual datasets, standardized benchmarks, prospective validation, and ethical governance.

Abstract

Background: Speech models (wav2vec 2.0, HuBERT, Whisper), large language models (GPT, LLaMA), and conversational AI have expanded computational speech analysis from handcrafted acoustic features to dialogue-based neurological assessment. How well these approaches address clinical practice has not been evaluated. Methods: We conducted a narrative review searching PubMed, Google Scholar, and IEEE Xplore, supplemented by Interspeech and ICASSP proceedings. Findings are organized along three layers: acoustic-motor (voice quality, prosody, articulation), language-transcript (lexical, syntactic, semantic, and discourse analysis), and integrated multimodal-conversational (interactive dialogue systems). Traditional acoustic biomarkers provide background; the primary focus is on foundation models, LLMs, and conversational AI. Findings: Speech foundation models outperform handcrafted features on several classification tasks but degrade on severely impaired speech due to domain mismatch with healthy training data. LLMs classify transcripts and score cognitive tests, but operate on text alone and cannot access acoustic-motor information. Conversational AI can administer cognitive screening through naturalistic dialogue, but validation is limited to small single-centre feasibility studies. Prospective clinical validation remains limited. Cross-linguistic generalizability is untested for most methods. Interpretation: The field is moving toward integrated speech-language assessment, but the gap between technical capability and clinical utility remains wide. Closing it requires diverse multilingual datasets, standardized benchmarks, prospective validation, and ethical governance.

Read PDF

Similar papers

Preprint Aug 2026

CASA: Content-Acoustic Speaking Assessment with Speech Encoder and Large Language Model

CASA, a simpler architecture combining Whisper-medium and Qwen3.5-2B that achieves state-of-the-art performance while providing a more interpretable separation between speech delivery and content, is proposed.

Nhan Phan, Ilona Lähteenmäki, Anna von Zansen et al. · 0 citations
Open access Aug 2026

Multimodal text-audio sentiment in clinical aphasia speech using NLP

Introduction Aphasia affects expressive and receptive communication and may influence the affective tone expressed during clinical speech tasks. This study presents an exploratory weakly supervised NLP analysis of positive/negative affective-tone proxies in AphasiaBank transcripts with paired audio. Methods We extracted sentence embeddings from DistilBERT (e∈R768) and recording-level acoustic summaries (MFCC13, ZCR, RMS, spectral centroid, and spectral bandwidth; a∈R17). Text and acoustic features were concatenated (x=[e;a]∈R785) and classified using Random Forest models. Sentiment labels were generated using an SST-2-derived weak-supervision pipeline and should be interpreted as pseudo-labels rather than clinical ground truth. To evaluate modality contribution and potential leakage, we compared text-only, audio-only, and fused text-audio models under utterance-level and recording-disjoint splits. A small five-rater evaluation was used to examine human judgment alignment. Results Under the recording-disjoint split, both the text-only and fused text-audio models achieved 97.9% accuracy and 0.791 macro-F1, while the audio-only model achieved 55.4% accuracy and 0.388 macro-F1. These results indicate that classification performance was primarily driven by textual embeddings, while the recording-level acoustic summaries did not improve performance over text-only features. Across the pseudo-labeled corpus, aphasic utterances were more often labeled negative than control utterances. Age-stratified summaries showed subgroup variation in pseudo-label distributions, but these patterns were treated descriptively because labels were model-derived. Human-rater agreement was low for aphasic utterances, indicating that affective-tone interpretation in fragmented clinical speech is ambiguous. Discussion These findings should be interpreted as exploratory evidence about weakly supervised affective-tone proxies, not as validated clinical sentiment recognition. The results highlight both the promise of clinical NLP for aphasia discourse analysis and the need for independent human-labeled validation, utterance-aligned acoustic features, and careful control of domain, task, age, and topic bias.

Shamiha Binta Manir, A. Kothari, William L. Gross et al. · 0 citations
Jul 2026

RW-Voice-EQ Bench: A Real World Benchmark for Evaluating Voice AI Systems

Current voice AI benchmarks typically evaluate isolated capabilities such as speech intelligibility, word error rate, or text-based dialogue quality, but they rarely test whether systems harness the acoustic information that distinguishes spoken language from its textual representation. To this end, we introduce the Real World Voice EQ Bench, a multidimensional benchmark for evaluating voice AI across text-to-speech (TTS), speech-to-speech (STS), speech understanding (SU), and automatic speech recognition (ASR). Our evaluations indicate that performance is highly dimension-specific. For TTS, naturalness, expressiveness, identity stability, and reliability are largely independent evaluation dimensions. For STS, access to audio does not guarantee use of vocal affect, and some agents remain largely transcript-driven. For SU, models perform unevenly across paralinguistic tasks. For ASR, real world accent, emotion, noise, and conversational conditions expose failures that are not captured by established clean-speech benchmarks. Together, these results show that voice AI should be evaluated as a profile of acoustic, expressive, interactional, and robustness capabilities rather than by a single aggregate score.

D. Ayllón, Alice Baird, Jeffrey Brooks et al. · 1 citation
Aug 2026

From acoustic features to clinical meaning: Emerging technologies for interpretable AI in pediatric speech

Interpretable artificial intelligence (AI) has become an increasingly important topic in speech communication as data-driven methods are used to support clinical decision-making in speech health. This talk highlights acoustic technologies that advance interpretable AI for pediatric preschool-age speech. Recent work on computer-assisted syllable analysis of continuous speech, published in the 2024 JASA special issue on Acoustic Cue-Based Perception and Production of Speech (Speights et al., 2024), demonstrates how linguistically structured, landmark-based acoustic representations improve robustness and interpretability relative to conventional spectral features and clinical metrics. Advances in automatic speech recognition further enable intelligibility to be characterized using Speech Intelligibility Probabilities (SIPs), which rely on model-intrinsic uncertainty rather than transcript-based accuracy metrics. Phoneme- and language-level recognition analyses additionally reveal age-stratified confusion structure associated with speech development. Together, these developments point toward future directions in which interpretable, developmentally grounded representations play a central role in advancing speech acoustics and clinical AI for pediatric speech.

Marisha L Speights, Vishal Shrivastava, Chethana Saligram · 0 citations
Preprint Aug 2026

Motor, Cognitive, or Corpus? What Survives Cross-Lingual Transfer in Speech-Based Parkinsons Disease Detection

A layer-wise analysis of nine SSL speech backbones using a low-capacity logistic regression probe reveals that the transferred discriminative signal lacks pathological specificity, highlighting critical limitations that must be addressed before speech-based pathology recognition models can be reliably deployed in clinical settings.

S. Kopar, Sam Gijsen, Abner Hernandez et al. · 0 citations
Jul 2026

Toward Generalizable Cognitive Impairment Detection with Speech-Based Multimodal Large Language Models

This work establishes a new state-of-the-art for CI identification, with the proposed method demonstrating superior cross-dataset generalization and the power of an LLM-based multimodal framework that fuses linguistic and acoustic data to enable robust, scalable, and non-invasive screening.

Ying-lei Huang, Xin Wang, Yuhan Su et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.