Correlation analysis indicates that acoustic similarity is the strongest predictor of fine-tuning performance, while phoneme inventory and typological similarity better explain zero-shot transfer.
Abstract
This paper investigates how language similarity can improve cross-lingual transfer for automatic speech recognition (ASR) in extremely low-resource settings. Warlpiri, an Australian Aboriginal language, has very limited transcribed speech data, making transfer learning essential. We propose a framework combining acoustic similarity from pre-trained speech models with linguistic similarity based on typology, phoneme inventories, grammatical, and syntactic features to rank high-resource source languages and evaluate their effectiveness for ASR transfer to Warlpiri. Experiments with Whisper show that acoustically and typologically similar languages outperform monolingual and multilingual baselines. Assamese and Hindi achieve substantial reductions in word and character error rates. Correlation analysis further indicates that acoustic similarity is the strongest predictor of fine-tuning performance, while phoneme inventory and typological similarity better explain zero-shot transfer.
In every setting, pre-adaptation on related auxiliary languages yields no practically meaningful improvements once as little as one hour of target-language data is available, suggesting that relatedness alone may not reliably predict transfer gains in large multilingual ASR, or constitute an effective strategy for extending such models to low-resource languages.
A. Florian, C. Amol, Hope Kerubo Ombaba et al.· 0 citations
Findings indicate that language-specific fine-tuning plays a more critical role than multilingual generalization in achieving accurate ASR for Indonesian and provide practical guidance for deploying ASR systems in low-resource language scenarios.
J. Hebert, Amalia Zahra· Bulletin of Electrical Engin...· 0 citations
SraVaani-1.0 achieves the lowest word error rate (WER) on a large number of language-dataset pairs while remaining competitive with the best-performing systems on high resource while being assessed exclusively on the VAANI benchmark.
Sujith Pulikodan, A. Basu, J. Pavankumar et al.· 1 citation· ⚡1
Automatic Speech Recognition (ASR) is typically evaluated using Word Error Rate (WER), which poorly reflects semantic similarity. While embedding-based metrics correlate better with human judgments, the respective roles of encoder and decoder-based Large Language Models (LLMs) remain underexplored. This paper presents a comparative study of both families for ASR evaluation. We analyze BERTScore and SemDist across different LLMs, layers, and pooling strategies, showing that both metrics can achieve strong correlation with human judgments when properly configured. For decoder models, we investigate generative LLMs in two settings: pairwise hypothesis selection via prompting and direct qualitative error classification. Our results show that encoder-based metrics remain highly competitive, while generative LLMs perform strongly in hypothesis comparison and improve the interpretability of ASR evaluation.
Thibault Bañeras-Roux, Shashi Kumar, Driss Khalil et al.· 0 citations
This work presents DonorRank, a learning-to-rank framework for predicting effective donor languages for zero-shot ASR and evaluates it on two multilingual speech corpora of Indic and African language families, finding transfer patterns that provide practical guidance for multilingual ASR in low-resource settings.
Akriti Dhasmana, Aarohi Srivastava, David Chiang· 0 citations
Forced aligners are an essential tool to facilitate access to the word and phone level of speech data, but they rely on acoustic models, pronunciation dictionaries, and complete word-level transcriptions that may not exist for low-resource languages. Alternatively, word-level alignments can be extracted from neural network-based tools for automatic speech recognition like Wav2Vec-BERT 2.0, avoiding the need to create bespoke acoustic models and dictionaries. This paper compares the ability of the Montreal Forced Aligner (MFA) and Wav2Vec-BERT 2.0 to produce word-level alignments for the low-resource variety Istro-Venetian, spoken in Croatia. Results indicate that MFA, which uses Gaussian mixture models with hidden Markov models, significantly outperforms Wav2Vec-BERT 2.0 on forced alignment. However, we show that Wav2Vec-BERT 2.0 can produce transcriptions with a word error rate of only 14.4%. Neural models can thus aid research on low-resource languages by creating transcriptions that can be used as input to forced aligners, and with sufficient data they may also perform acceptably on forced alignment.
Austin Jones, Massimo Daul, Margaret E. L. Renwick et al.· Journal of the Acoustical So...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.