Skip to content
Preprint

Analyzing Speech Condition Effects in Dysarthric ASR: A Layer-wise Probing Study

Aug 2026 · 0 citations · 21 references
Computer Science

TL;DR

A layer-wise probing analysis of a transformer ASR encoder on Mandarin dysarthric speech under three transcript-matched conditions links representation analysis to parameter-efficient fine-tuning and motivate layer-aware adaptation for low-resource Mandarin dysarthric ASR.

Abstract

Automatic speech recognition (ASR) performance degrades sharply on dysarthric speech, yet how disordered articulation reshapes a model's internal representations is underexplored. We conduct a layer-wise probing analysis of a transformer ASR encoder on Mandarin dysarthric speech under three transcript-matched conditions: original dysarthric speech, speaker conditioned zero-shot TTS resynthesis, and unconditioned TTS. Probing reveals a task- and condition-dependent representation hierarchy: phoneme boundary information remains weak across all layers for dysarthric speech; phoneme identity is recoverable in deep layers for synthetic speech, but remains poor for dysarthric speech; and recognition difficulty is concentrated in the deepest layers. Furthermore, lexical tone is a persistent error source across all conditions. Guided by these insights, layer-selective LoRA shows that mid-layer adaptation (layer 7 or layers 5-8) recovers near-full encoder performance on dysarthric speech within 6.67% and 2.89% relative margins while training only 0.16% and 0.65% of adapter parameters. Conversely, upper-layer adaptation benefits synthetic speech more than dysarthric speech. These findings link representation analysis to parameter-efficient fine-tuning and motivate layer-aware adaptation for low-resource Mandarin dysarthric ASR.

View source

Similar papers

Open access Jul 2026

Acoustic and Perceptual Characteristics of Compressed Dysarthric Speech: Implications for Teleassessment.

PURPOSE This study aimed to evaluate how speech signal compression algorithms affect the acoustic and perceptual characteristics of dysarthric speech. As telepractice becomes more common in speech-language pathology, particularly in the assessment of dysarthria, understanding the impact of telecommunication compression on speech fidelity is essential for clinical decision making. METHOD Speech samples were collected from 15 individuals with dysarthria, recorded locally and simultaneously via Zoom (Version 6.2.11). The Opus codec was used to process the locally recorded speech samples under fullband, wideband, and narrowband compression conditions. Acoustic measures were derived for all samples. Twenty experienced speech-language pathologists (SLPs) rated vocal quality and nasality for sustained vowels and rated articulatory precision and orthographically transcribed sentences across each condition. RESULTS Narrowband compression was associated with significantly degraded voice quality measures (e.g., harmonic-to-noise ratio, shimmer) and reduced transcription accuracy. Despite these changes in acoustics and intelligibility, ratings of articulatory precision, nasality, and vocal quality remained stable across conditions, suggesting SLPs could adjust for compression artifacts when making perceptual judgments. CONCLUSIONS Speech compression, particularly in low-bandwidth (i.e., narrow) condition, impacts acoustic fidelity and intelligibility, though expert listeners maintain reliable perceptual judgments. These findings underscore the need to consider compression effects in telepractice and highlight the importance of developing optimized protocols for remote dysarthria assessment.

Kelvin Tran, Rene L. Utianski, V. Berisha et al. · 0 citations
Jul 2026

Re-Sonance: A Dysarthric Asynchronous Real-Time Speech Conversion System Based on a Three-Stage Cascaded ASR-LLM-TTS Architecture

Individuals with dysarthria face significant challenges in professional speaking scenarios such as conferences, presentations, and meetings, where real-time communication is crucial. While existing Augmentative and Alternative Communication (AAC) systems provide basic support, they often fail to meet the demands of professional speaking environments due to high latency and unnatural speech patterns. This paper presents Re-Sonance, a novel LLM-enhanced speech-driven AAC system designed for real-time professional speaking scenarios. By integrating Whisper ASR, Qwen LLM, and CosyVoice TTS, Re-Sonance achieves improved speech intelligibility and naturalness while maintaining real-time performance. Both subjective and objective evaluations using a Mandarin dysarthric speech dataset demonstrate that our speech reconstruction approach significantly improved intelligibility while preserving semantic coherence for speakers with mild to moderate dysarthria. Although performance remains limited for severe dysarthria cases, our findings validate the potential of LLM-based methods for enhancing speech-driven AAC systems, paving the way for more effective and accessible communication technologies.

Yuxuan Wu, Yifan Xu, Jun-Kun Wang et al. · 0 citations
#natural language process... Preprint Sep 2026

Choosing a PEFT Variant for Per-Patient Dysarthric ASR: A Single-Speaker Case Study on Two ASR Bases

Per-patient adapters are the preferred production architecture for dysarthric automatic speech recognition (ASR), yet parameter-efficient fine-tuning (PEFT) variants have not been compared in the speaker-dependent, per-patient regime. We present a single-speaker case study comparing seven LoRA-family methods (LoRA, QLoRA, AdaLoRA, DoRA, LoHA, VeRA, VB-LoRA) on two production bases (Whisper-large-v3 with Hungarian fine-tuning, and a multilingual Qwen3-ASR-1.7B checkpoint) for one post-stroke Hungarian male speaker (S1, 409 utterances; severe dysarthria on auditory-perceptual clinical assessment). Attention-projection adapters substantially improve CER on both bases. Across three seeds, a paired bootstrap detects no significant LoRA-DoRA difference (p>0.5; 13.86/13.90 % CER on Whisper, 28.10/28.33 % on Qwen3-ASR), so we adopt the simpler, cheaper LoRA. Real 4-bit (NF4) QLoRA is worse on every seed and both bases (14.56/30.09 % CER) with no memory saving at this scale, and LoHA, VeRA, VB-LoRA and AdaLoRA do not reach the LoRA family, though LoHA still gives an 18.6 % relative CER reduction on Whisper. On the same base, full fine-tuning is more accurate (11.43 % CER), but a 115 MB LoRA that also adapts the feed-forward blocks reaches within 0.66 pp of it at approximately 3.7 % of the per-patient storage. A 6-point enrollment grid shows about 5 min of patient audio captures 45.6 % of the zero-shot-to-30-min CER reduction, with further gains at 10 and 30 min (caveat: one speaker, one language, severe post-stroke dysarthria). Training scripts and recipes will be released, source-available under a research-use licence, on publication.

B. Muller, László Tóth, L. Roberts · 0 citations
Aug 2026

Leveraging automatic forced alignment of reading passages in dysarthria

Manual acoustic-phonetic segmentation of dysarthric speech is challenging due to the degraded nature of the acoustic signal. The Montreal Forced Aligner (MFA) is a gold-standard approach for automatic segmentation in non-disordered speech, but its efficacy for speakers with dysarthria has not been evaluated. We evaluate the MFA’s performance on vowel segmentation in a standardized reading passage produced by five talkers with dysarthria and five healthy controls. MFA input includes speech audio and an orthographic transcript and uses pre-trained acoustic models and pronunciation dictionaries to time-align word and phone boundaries. We manipulate three alignment conditions that differ in the amount of pre-processing a researcher may choose to perform on their input transcripts, namely (1) the full text of the reading passage, (2) the passage segmented into utterances, and (3) the passage segmented into utterances, with speech errors corrected in the text. We compare the force aligned output with manually segmented vowels as a function of condition, speaker group, and vowel. Alignment across all conditions was high ( > 90% accuracy) for controls, but varied widely for dysarthric speech. Alignment accuracy for dysarthric speakers ranged from 20% (full-text) to 76.4% (segmented-corrected-text). Findings will guide best-practices for leveraging automatic acoustic tools in disordered speech research.

Thea Knowles, Maura Philippone, Maria Cuervo Cano et al. · 0 citations
Open access Sep 2026

Adaptive Attention and Pyramid Pooling for High-Performance Dysarthric Speech Recognition

Speech pathology is a multidisciplinary field that assesses, diagnoses, and manages communication disorders caused by neurological, developmental, structural, and degenerative conditions. Among these disorders, dysarthria is a motor speech impairment characterized by reduced speech intelligibility and abnormal articulatory patterns resulting from neuromuscular dysfunction. Automatic dysarthria assessment remains challenging because of substantial inter-speaker variability, diverse severity levels, and the complex temporal–spectral characteristics of pathological speech signals. To address these challenges, this paper proposes an attention-enhanced convolutional neural network for automatic dysarthria detection and severity classification. The proposed architecture integrates a convolutional block attention module to enhance channel-wise and spatial feature representations, a coordinate attention module to capture long-range contextual dependencies while preserving positional information, and spatial pyramid pooling to extract robust multi-scale representations and accommodate input variability. Hierarchical convolutional layers with rectified linear unit activation and max-pooling operations progressively learn discriminative speech representations, followed by fully connected layers and a softmax classifier for prediction. The proposed framework was comprehensively evaluated on two publicly available datasets: TORGO Dataset 1, containing dysarthric and healthy control speech recordings, and TORGO Severity Dataset 2, comprising four dysarthria severity classes. Mel-frequency cepstral coefficients, together with their first- and second-order temporal derivatives and complementary acoustic descriptors, were extracted to characterize the speech signals. Extensive experiments, including cross-validation, repeated runs, ablation studies, Shapley additive explanations-based interpretability analysis, and statistical significance analysis, demonstrate the effectiveness and robustness of the proposed framework. On TORGO Dataset 1, the proposed model achieved 95.00% accuracy, 94.50% precision, 95.25% recall, and a 94.75% F1-score, outperforming conventional machine learning and deep learning baselines. On TORGO Severity Dataset 2, the proposed framework achieved 99.10% classification accuracy, an F1-score of 0.9910, and an area under the curve of 0.9999, demonstrating strong performance in multi-class dysarthria severity classification. These findings highlight the proposed model’s strong discriminative capability, robustness, and interpretability, indicating its potential as a reliable computer-aided decision-support system for dysarthria detection and severity assessment.

Z. Tarek, Tarek Abd El-Hafeez, Esraa Hassan et al. · 0 citations
Open access 2026

LoRA-MoE Fine-Tuning for Improved Speech Recognition in People With Parkinson’s Disease

LoRA-MoE, a parameter-efficient adaptation method that combines Low-Rank Adaptation with a mixture of experts (MoE) to improve speech recognition for individuals with Parkinson’s disease (PD), demonstrates consistent improvements and stable performance across all severity levels, and its performance is robust to the number of experts.

Seojin Yoon, Ryul Kim, Sangmin Lee · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.