Skip to content
Open access

DA-ICL: Distribution-Aware In-Context Learning for Arabic Automatic Speech Recognition Error Correction

2026 · IEEE Access · Vol 14, pp. 106917-106926 · 0 citations · 29 references
Computer Science

TL;DR

A novel two-stage framework for accurate and efficient Arabic ASR enhancement, combining HuBERT-based acoustic modeling with LLM-based DA-ICL correction and LoRA-efficient adaptation yields a robust, accurate, and scalable solution for Arabic ASR, effectively bridging the gap between acoustic signal and linguistic knowledge.

Abstract

Arabic Speech Recognition (ASR) faces compounded challenges due to rich dialectal variation, morphological complexity, and data scarcity. While self-supervised speech models such as HuBERT excel in acoustic representation, they lack the deep linguistic reasoning needed to resolve ambiguities unique to Arabic. Large Language Models (LLMs) offer complementary grammatical and semantic knowledge, yet their role in systematic, real-time Arabic ASR error correction remains underexplored. In this work, we propose a novel two-stage framework for accurate and efficient Arabic ASR enhancement. First, we fine-tune a HuBERT model on the Common Voice Arabic corpus, establishing a baseline word error rate (WER) of 19.3%. Second, we introduce Distribution-Aware In-Context Learning (DA-ICL), a prompting strategy that supplies the Arabic LLM Aya-23-8B with a curated set of few-shot examples derived from a systematic taxonomy of ASR error types, including phonetic confusions and morpho-orthographic errors. DA-ICL enables precise, structurally faithful corrections, reducing WER to 9.6% without undesirable sentence rephrasing. To address domain shift and catastrophic forgetting, we further apply Low-Rank Adaptation (LoRA) to adapt a pre-trained HuBERT model to new domains parameter-efficiently. This approach reduces out-of-domain WER from 67% to 24% while preserving in-domain performance, demonstrating improved generalization without full fine-tuning. Our results confirm that combining HuBERT-based acoustic modeling with LLM-based DA-ICL correction and LoRA-efficient adaptation yields a robust, accurate, and scalable solution for Arabic ASR, effectively bridging the gap between acoustic signal and linguistic knowledge. Our framework achieves a WER of 9.6% on Common Voice Arabic, significantly outperforming Whisper-large (47.49% zero-shot, 37.89% with LoRA fine-tuning) and demonstrating the effectiveness of our linguistically-aware approach for Arabic speech recognition.

Read PDF

Similar papers

Preprint Aug 2026

SraVaani 1.0: Scaling Inclusive Speech Recognition for Indic Languages

SraVaani-1.0 achieves the lowest word error rate (WER) on a large number of language-dataset pairs while remaining competitive with the best-performing systems on high resource while being assessed exclusively on the VAANI benchmark.

Sujith Pulikodan, A. Basu, J. Pavankumar et al. · 1 citation · ⚡1
Preprint Aug 2026

AraSSM: A bidirectional state-space encoder for Arabic masked language modeling

A bidirectional Mamba encoder pretrained via masked language modeling on a corpus combining Arabic Wikipedia and CulturaX text is introduced, trained end-to-end on four consumer-grade NVIDIA RTX 2080Ti GPUs (11GB) over approximately ten days.

Ahmed Amine Aliane, H. Aliane, N. Semmar · 0 citations
Preprint Aug 2026

MERaLiON-GR: Speech Gender Recognition Model for English and SEA Languages

We present MERaLiON-GR, a speech gender recognition system that performs binary classification (female / male) on English and Southeast Asian (SEA) languages. The model finetunes MERaLiON-SpeechEncoder-2, a large conformer based transformer pre-trained on a broad speech corpus, and applies parameter efficient fine-tuning via Low-Rank Adaptation (LoRA) to adapt the encoder to the gender recognition task, and appends a multi-scale ECAPA-TDNN down stream network with attention pooling and a lightweight linear classifier. Extensive evaluations across multilingual Singaporean and Southeast Asian languages (English, Chinese, Malay, Tamil, Thai, Vietnamese, Indonesian, and Khmer) show that MERaLiON-GR consistently surpasses the state-of-the-art gender recognition model Vox-Profile and a large Audio-LLM, in both full-utterance and segment level evaluation modes. The results underscore the value of dedicated speech models in achieving accurate paralinguistic understanding and strong cross-lingual generalization.

Qiongqiong Wang, A. Aw, Nancy F. Chen et al. · 0 citations
Open access Jul 2026

Explainable AI Framework for Cognitive and Pragmatic Analysis of Classical Arabic Narratives Using Large Language Models

The results demonstrate that linear discriminative models and appropriate lexical feature engineering can provide a very accurate and interpretable baseline for the development of natural language processing algorithms for Arabic.

Hamood Mohammed Alrumhi, Muhammad Asshad, Amjed Abbas Ahmed et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.