Skip to content

Efficient Adaptation of English Language Models for Morphologically Rich and Underrepresented Languages: The Case of Arabic

2026 · International Conference on Language Resources and Evaluation · pp. 10485-10496 · 0 citations · 44 references
Computer Science

TL;DR

A resource-efficient adaptation of the English-pretrained ModernBERT for Arabic, employing continued pretraining on large Arabic corpora followed by lightweight head-only fine-tuning with a frozen encoder, demonstrating that modern English encoder architectures can be efficiently transferred to Arabic through language-adaptive pretraining.

View source

Similar papers

Preprint Aug 2026

Efficient Multilingual Neural Machine Translation via Corpus-Driven Vocabulary Pruning: An English-Arabic Case Study

This paper proposes a general optimization framework that combines a vocabulary pruning method with a targeted fine-tuning protocol for MNMT models, and reduces the vocabulary size from over 128,000 to approximately 10,000 tokens, enabling a 60% memory saving without any loss in performance.

A. A. Aliane, N. Semmar, H. Aliane · 0 citations
Open access Sep 2026

A Supervised Morphosyntactic Analyzer for Modern and Classical Arabic Text

Arabic morphological analysis remains a central challenge for natural language processing (NLP). This challenge is embedded within the complex, rich, and multi-layered morphology of Arabic, where orthographic words encode enclitics, root-and-pattern derivations, and other morphosyntactic features. Another dimension of...

M. Sawalha · 0 citations
Preprint Aug 2026

AraSSM: A bidirectional state-space encoder for Arabic masked language modeling

A bidirectional Mamba encoder pretrained via masked language modeling on a corpus combining Arabic Wikipedia and CulturaX text is introduced, trained end-to-end on four consumer-grade NVIDIA RTX 2080Ti GPUs (11GB) over approximately ten days.

A. A. Aliane, H. Aliane, N. Semmar · 0 citations
Aug 2026

STAR: instruction tuning for Arabic across tasks, datasets, and models

An in-depth evaluation of instruction tuning for Arabic NLP tasks using three prominent LLMs: LLaMA 3.1-8B, AceGPT-v2-8B, and Qwen3-8B shows that instruction tuning consistently improves performance across most tasks, with notable variations in effectiveness across different tasks and prompts.

Maged Saeed Al-shaibani, Zaid Alyafeai, Irfan Ahmad · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.