Skip to content

From scalar confidence to structured uncertainty: Modeling child speech intelligibility from acoustic model evidence

Aug 2026 · Journal of the Acoustical Society of America · Vol 159, pp. A284-A284 · 0 citations

TL;DR

A Speech Intelligibility Probability (SIP) framework is proposed that conceptualizes intelligibility as acoustically induced model uncertainty rather than as word error or decoder confidence, and indicates that internal model uncertainty serves as a robust proxy for speech intelligibility.

Abstract

Accurate assessment of children’s speech intelligibility is critical for identifying speech sound disorders. Clinical practice relies on human transcription and perceptual ratings that are costly, inconsistent, and difficult to scale. Existing ASR-based intelligibility metrics offer scalability but face two fundamental issues: (i) training-data mismatch, as adult-centric models produce unstable representations for developmentally variable child speech, and (ii) decoding artifacts, e.g., softmax sharpening, that yield spuriously confident scores and mask acoustic ambiguity. We propose a Speech Intelligibility Probability (SIP) framework that conceptualizes intelligibility as acoustically induced model uncertainty rather than as word error or decoder confidence. Using the Speech Exemplar and Evaluation Database, we extracted token-level posterior trajectories from Whisper Large-v2 under controlled decoding and analyzed them using a multidimensional uncertainty profile capturing distributional dispersion (Tsallis entropy) and temporal instability (coherence), thereby characterizing acoustic-driven uncertainty beyond post-decoding confidence scores tied to committed hypotheses. We compared SIP against Hustad et al. (2021) developmental growth curves. It aligned with human perceptual difficulty: typically developing speech clustered near mid-range expectations, while disordered speech concentrated in the bottom decile. This correspondence indicates that internal model uncertainty serves as a robust proxy for speech intelligibility.

View source

Similar papers

#machine learning Preprint Oct 2026

Beyond Decodability: Do Acoustic Factors Drive Predictions in Speech-Based Alzheimer's Assessment?

Speech-based Alzheimer's disease (AD) assessments increasingly rely on pretrained self-supervised learning (SSL) models that learn acoustic representations directly from raw audio, exposing the model to recording factors. We ask whether such factors are merely encoded in SSL representations or can systematically alter...

Serli Kopar, Alkis Koudounas, R. Rane et al. · 0 citations
#artificial intelligence Preprint Sep 2026

SpeechCritic: Learning a Diagnostic Speech Judge from Limited Human Preferences

SpeechCritic is introduced, which learns a diagnostic judge in a reference-conditioned cross-lingual setting from only about 300 human-labeled comparisons, and shows that the pipeline is language-pair agnostic by instantiating it on both English-Japanese and English-Spanish.

Ming-Yue Huo, Shivam Mehta, Bhavin Jawade et al. · 0 citations
Preprint Aug 2026

CASA: Content-Acoustic Speaking Assessment with Speech Encoder and Large Language Model

CASA, a simpler architecture combining Whisper-medium and Qwen3.5-2B that achieves state-of-the-art performance while providing a more interpretable separation between speech delivery and content, is proposed.

Nhan Phan, Ilona Lähteenmäki, Anna von Zansen et al. · 0 citations
Open access Sep 2026

Toolkit for acoustic–phonetic analysis of naturalistic speech data

This work demonstrates TAPA on the 2016 U.S. presidential debate, and suggests that TAPA can be used to increase access to naturalistic speech data and speed up the processing timeline with experts' supervision.

Ethan Kutlu, Emerson Peters, Ciara Tapanes et al. · 0 citations
Preprint Sep 2026

ART-NAD: An Articulatory Inversion-based Neural Acoustic Distance for Pathological Speech Intelligibility Assessment

Speech assessment tools for speakers with speech pathology must be both accurate and interpretable if they are to be adopted in clinical practice. Existing reference-audio measures such as the Neural Acoustic Distance (NAD) reach high speaker-level correlations with listener intelligibility scores but operate on self-s...

B. Halpern, Thomas B. Tienkamp, D. Abur et al. · 0 citations
Aug 2026

Interpretable but Not Necessarily Meaningful: Language, Voice, and Clinical Inference in AI-Based Mental Health Prediction-A Commentary on Wang and Sambamoorthi (2026).

An interpretable and fairness-aware Random Forest model that predicts mental health disorders using acoustic features derived from the Bridge2AI-Voice dataset is developed, and the demonstration of strong gender fairness represents a commendable step toward responsible artificial intelligence in voice science.

Ali Khodi · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.