Skip to content

Author

J. Pushparaj

2 papers indexed here

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

#small language model Open access Sep 2026

Exploring Wav2Vec2 XLSR-53 and Whisper-small for Telugu language speech-to-text: a fine-tuning approach

Online voice-based applications and speech communication have grown as a result of the revolutionary rise of smart gadgets and social media. The rapid advancement of deep learning (DL) has transformed the field of audio processing, enabling smooth human-computer interaction. DL approaches have been used to develop speech-to-text (STT) systems across various languages and topics. These models require a large amount of training data: extensive corpora of continuous speech utterances collected from numerous speakers, along with their corresponding transcripts. In this study, we explore the use of state-of-the-art pre-trained models like Wav2Vec2 XLSR-53 and Whisper-small for developing STT systems in the Telugu language, addressing the challenge of limited data availability and demonstrating satisfactory results [23.62% Word Error Rate (WER), 4.12% Character Error Rate (CER)] even when fine-tuned on a smaller dataset. To evaluate model performance, we employed a k-fold cross-validation approach with values of k = 2 to k = 5, and compared the results with the conventional train-test split method. The results indicate that at k = 5, the Wav2Vec2 XLSR-53 model achieved a cross-validation Character Error Rate(WER) of 23.62% and a Character Error Rate (CER) of 4.12%, and the Whisper-small model yielded a cross-validation WER of 28.67% and a CER of 5.48%. These findings suggest that the k-fold cross-validation strategy, at k = 5, enhances the robustness of Wav2Vec2 XLSR-53 in low-resource language scenarios where training data is limited. For Whisper-small, however, the baseline train-test split (WER: 27.76%) outperformed all k-fold configurations tested (k = 2: 32.05%, k = 3: 29.72%, k = 4: 29.40%, k = 5: 28.67%), indicating that the benefit of k-fold cross-validation observed for Wav2Vec2 XLSR-53 does not generalize across model architectures. Additionally, the models were evaluated on unseen noisy data, and both models demonstrated satisfactory performance.

Muzaffar Ahmad Dar, Sri Gani Kaarthikeya Kammula, J. Pushparaj et al. · 0 citations
Open access Jul 2026

A comparative analysis of pretrained Wav2Vec XLSR-53 and Whisper-Small models for automatic speech recognition in the Telugu language

This study presents the development of an automatic speech recognition (ASR) system tailored for Telugu, one of the widely spoken Indian languages. In recent years, deep learning (DL) techniques have been applied to develop ASR systems across various languages and domains. These models, however, require substantial training resources and extensive corpora of continuous speech composed from multiple dialectal speakers, along with their corresponding transcripts. This paper investigates the effectiveness of pre-trained models like Wav2Vec XLSR-53 and Whisper-Small for developing ASR systems for the Telugu language, addressing the challenge of limited data availability and demonstrating satisfactory results even when fine-tuned on a smaller dataset. We utilized approximately 20 h of speech data comprising 17,421 sentences of the Telugu language. The models are fine-tuned on four publicly available datasets, including OpenSLR, Common Voice, IndicVoices, and IndicTTS, to introduce greater diversity in both speaker demographics and linguistic content. The Wav2Vec XLSR-53 model achieved a word error rate (WER) of 27.3% and a character error rate (CER) of 6.8% on the test dataset, whereas the Whisper-Small attained a WER of 28.67% and a CER of 7.55%. In addition, performance of the models was evaluated by introducing noise to both individual datasets as well as a combined noise dataset. The results show that, on the combined noise dataset, Wav2Vec XLSR-53 achieved a WER of 19.59% and a CER of 4.58%, while Whisper Small obtained a lower WER of 13.97% and a CER of 3.54%. These results underscore the usefulness of leveraging pre-trained architectures in low-resource linguistic scenarios such as Telugu.

J. Pushparaj, Muzaffar Ahmad Dar, Sri Gani Kaarthikeya Kammula et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.