Foundations for pediatric vocal biomarkers: age-aware phoneme recognition and latent-space error analysis
Abstract
Introduction Pediatric speech sound disorders (SSDs) affect many young children and are commonly assessed through auditory-perceptual judgments and IPA transcription, which are limited by listener bias, variable interrater reliability, and difficulty attributing deviations at the phoneme level. We present a developmentally informed framework for pediatric vocal biomarker foundations that prioritizes age-aware, phoneme-resolved interpretability over utterance-level accuracy alone. Methods Using a CAAP-derived subset of the SEED corpus (27–94 months) with clinically specified phoneme targets and SSD labels, we construct age-stratified phoneme profiles and quantify error structure via phoneme error rate (PER) and age-banded confusion signatures. We operationalize co-articulation as transition dynamics from MFCC trajectories and their first and second derivatives, reflecting the velocity and acceleration of spectral change across adjacent segments. We learn compact acoustic representations with a variational autoencoder (VAE) and model temporal evolution with BiLSTMs, including attention, to characterize disorder-relevant instability in latent trajectories. For phoneme transcription, we train BiLSTM-CTC sequence models on clinically elicited speech and evaluate disorder classification and phoneme substitution patterns within age groups, then stress-test generalization on ECSC "Frog Story" narratives from TalkBank/CHILDES. Results Co-articulation transition magnitudes varied systematically by age and clinical group. Attention improved temporal classification specifically in younger children, and PER dropped markedly between ages 3–4 and 4–5. On long-form naturalistic narratives, CTC training and inference remained numerically stable, but greedy decoding produced degenerate output dominated by blank and repetition tokens. Discussion Together, these results support an age-aware approach to pediatric phoneme analytics, in which articulatory coordination, not just phoneme identity, carries developmentally and clinically relevant information; and phoneme-level performance should be interpreted relative to developmental stage rather than a single fixed benchmark. The decoding failures on naturalistic speech indicate a bottleneck specific to decoding rather than to the underlying representations, motivating constrained, hierarchical, or duration-aware decoding as a tractable next step. These findings position age-stratified, phoneme-resolved analysis as a foundation for interpretable and scalable pediatric screening and biomarker development.