Adaptive Attention and Pyramid Pooling for High-Performance Dysarthric Speech Recognition
Speech pathology is a multidisciplinary field that assesses, diagnoses, and manages communication disorders caused by neurological, developmental, structural, and degenerative conditions. Among these disorders, dysarthria is a motor speech impairment characterized by reduced speech intelligibility and abnormal articulatory patterns resulting from neuromuscular dysfunction. Automatic dysarthria assessment remains challenging because of substantial inter-speaker variability, diverse severity levels, and the complex temporal–spectral characteristics of pathological speech signals. To address these challenges, this paper proposes an attention-enhanced convolutional neural network for automatic dysarthria detection and severity classification. The proposed architecture integrates a convolutional block attention module to enhance channel-wise and spatial feature representations, a coordinate attention module to capture long-range contextual dependencies while preserving positional information, and spatial pyramid pooling to extract robust multi-scale representations and accommodate input variability. Hierarchical convolutional layers with rectified linear unit activation and max-pooling operations progressively learn discriminative speech representations, followed by fully connected layers and a softmax classifier for prediction. The proposed framework was comprehensively evaluated on two publicly available datasets: TORGO Dataset 1, containing dysarthric and healthy control speech recordings, and TORGO Severity Dataset 2, comprising four dysarthria severity classes. Mel-frequency cepstral coefficients, together with their first- and second-order temporal derivatives and complementary acoustic descriptors, were extracted to characterize the speech signals. Extensive experiments, including cross-validation, repeated runs, ablation studies, Shapley additive explanations-based interpretability analysis, and statistical significance analysis, demonstrate the effectiveness and robustness of the proposed framework. On TORGO Dataset 1, the proposed model achieved 95.00% accuracy, 94.50% precision, 95.25% recall, and a 94.75% F1-score, outperforming conventional machine learning and deep learning baselines. On TORGO Severity Dataset 2, the proposed framework achieved 99.10% classification accuracy, an F1-score of 0.9910, and an area under the curve of 0.9999, demonstrating strong performance in multi-class dysarthria severity classification. These findings highlight the proposed model’s strong discriminative capability, robustness, and interpretability, indicating its potential as a reliable computer-aided decision-support system for dysarthria detection and severity assessment.