Diagnostics of speech disorders in psychological and speech therapy practice: crosslingual validation of the SlowFast model
Abstract
In the authors’ previous study, the SlowFast two-stream architecture demonstrated high efficiency in diagnosing dyslalia on Russian-language data (98.0% accuracy on real clinical recordings). However, the question of the model's cross-lingual robustness remained open. This paper evaluates the ability of a SlowFast model trained on Russian to correctly classify speech disorders in Polish-speaking children. The open PAVSig dataset (N=201 children, 66,781 audio segments with double expert diagnosis of sigmatism) was used as the target corpus. In zero-shot evaluation mode (without any retraining on Polish data), the without the need to collect large labeled corpora for each model achieved 87.3% accuracy, 86.8% precision, 88.1% recall, 87.4% F1-score, and 0.92 AUC-ROC. Minimal adaptation (finetuning on 10% of the data) increased accuracy to 91.8%, and on 25% of the data – to 94.2%. Error analysis revealed that the main difficulties are associated with acoustic differences in Polish sibilant sounds (sz, cz, ż), accounting for 58% of all errors in zero-shot mode, as well as background noise (27%) and short segment duration (15%). After fine-tuning on 25% of the data, the proportion of errors due to cross-lingual acoustic differences decreased to 22%. The results confirm the high cross-lingual transferability of the SlowFast architecture and open prospects for creating multilingual automated speech disorder diagnostic systems new language.