Skip to content
Conference

An Integrated Semi-Supervised Framework for Low-Resource Language ASR: Acoustic Enhancement, LoRA Fine-Tuning, and Hybrid Post-Processing for Amazigh Variants

Aug 2026 · 2026 6th International Conference on Emerging Smart Technologies and Applications (eSmarTA) · pp. 1-8 · 0 citations · 34 references

Abstract

Low-resource languages continue to face significant challenges in automatic speech recognition (ASR), especially those with limited annotated corpora and significant variant variation. The Amazigh language family, which is spoken throughout North Africa, has very little digital infrastructure and is severely underfunded. In this paper, we present a multi- variant Amazigh ASR system that integrates linguistically-informed post-processing correction and audio enhancement preprocessing with refined multilingual self-supervised models is presented. We address three key issues: (1) acoustic degradation in field recordings; (2) substantial inter-variant variation; and (3) extreme data scarcity. Our enhanced pipeline adds a semi-supervised learning framework that includes: (1) data augmentation through synthetic speech generation and aggressive spectral perturbation, which increases the training corpus to 15 hours; (2) self-training on unlabeled data using pseudo-labeling; (3) VoiceFixer-based audio restoration; and (4) hybrid n-gram/Levenshtein post-correction. WER of 18.7% ± 2.3% (95% CI), CER of 9.4% ± 1.6%, and PER of 12.1% ± 1.9% show statistically significant improvements, with a relative WER reduction of 34.2% compared to baseline Whisper (p < 0.001). Component contributions are quantified by ablation studies: audio preprocessing results in a relative improvement of 8.3%, while post-processing adds 12.7% reduction. In addition to offering a broadly applicable framework for the preservation of low-resource languages, this work sets new state-of-the-art for Amazigh ASR.

View source

Similar papers

Preprint Aug 2026

Breaking the Curse of Multilinguality in Many-to-Many Speech-to-Text Translation via a Resource-Aware Mixture of Speech Encoders

Empirical analyses show that MoSE improves high-, medium-, and low-resource languages simultaneously, with the largest gains on low-resource speech, thereby breaking the curse of multilinguality without compromising high-resource performance.

Yexing Du, Kaiyuan Liu, You-Cheng Pan et al. · 0 citations

Models are Zero-Shot Text

Experimental results show that VALL-E outperforms the state-of-the-art zero-shot TTS system in terms of speech naturalness and speaker similarity and could preserve the speaker’s emotion and acoustic environment from the prompt in synthesis.

Unknown authors · 0 citations
#natural language process... Preprint Sep 2026

SEAR: Segment-Evidence-Aware Routing for Weak-to-Strong Multilingual Speech MCQ

This paper describes our system for Task~2 of the second Multilingual Conversational Speech Language Model (MLC-SLM) Challenge. We adapt Qwen3-Omni-30B-A3B-Instruct with a segment-evidence-aware data and post-training pipeline. A language model converts timestamped ASR into coherent event spans, which are expanded by a...

H. Le, L. Nguyen, Minh Tri Dao · 1 citation
#natural language process... Preprint Sep 2026

Qwen-Audio-3.0-ASR Technical Report

In recent years, automatic speech recognition (ASR) has witnessed transformative advancements driven by three complementary paradigms: data scaling, model scaling, and deep integration with large language models (LLMs). However, bridging the gap between academic benchmark performance and real-world production utility r...

Chuan-Meng Bian, Da-Ren Chen, Pei-Xin Chen et al. · 2 citations
Jul 2026

DONDO: Open w2v-BERT Speech-Recognition Base Models for African Languages

We present DONDO, a family of open, permissively licensed automatic speech recognition (ASR) base models for African languages, built on the w2v-BERT 2.0 self-supervised speech encoder. DONDO comprises twenty-one monolingual models and five multilingual models spanning twenty-seven language varieties across Ghana, Sier...

P. Azunre, Naafi Dasana Ibrahim, Joel Budu et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.