#artificial intelligence
May 2026
When Audio-Language Models Fail to Leverage Multimodal Context for Dysarthric Speech Recognition
A benchmark built on the Speech Accessibility Project (SAP) dataset is introduced that tests whether diagnosis labels, clinician-derived speech ratings, and progressively richer clinical descriptions improve transcription accuracy for dysarthric speech, finding that current models do not meaningfully use this context.
P. Moure, Niclas Pokel, Bilal Bounajma et al.
· arXiv.org · 2 citations