Skip to content
Preprint

Biomedical Machine Translation for Low-Resource Arabic-Script Languages via Cross-Lingual Transfer and LoRA Adapter Merging

Jul 2026 · 0 citations · 51 references
Computer Science

TL;DR

A systematic study of healthcare-domain cross-lingual transfer to address the scarcity of biomedical NMT resources for Arabic-script languages, finding LoRA adapter merging works surprisingly well for closely related languages, even without target-language biomedical data.

Abstract

We present a systematic study of healthcare-domain cross-lingual transfer to address the scarcity of biomedical NMT resources for Arabic-script languages. We use Arabic and Persian as higher-resource pivots to improve translation for \textbf{four severely low-resource} targets: Dari (Afghan Persian, a standardised variety of Persian), Pashto, Sorani Kurdish (Central Kurdish, a major standardized variety of Kurdish), and Urdu (closely related to Hindi). Using LoRA fine-tuning on small decoder-only LLMs, we train \textit{domain-specific pivot adapters} and evaluate \textbf{three transfer strategies}: few-shot in-context learning, minimal supervised adaptation, and, to the best of our knowledge, for the first time in this setting, zero-data LoRA adapter merging. Supervised adaptation with just 500 sentences achieves near pivot-language quality for Dari (CHrF++ 41.01) and meaningful gains for Urdu (28.88), while adapter merging reaches within 3.5 CHrF++ of supervised adaptation for Dari at zero additional cost. Pashto and Sorani Kurdish remain insufficient for high-stakes clinical deployment exposing the limits of cross-lingual transfer when structural distance from the pivots is too great. LoRA adapter merging works surprisingly well for closely related languages, even without target-language biomedical data.

View source

Similar papers

Open access 2026

Using English-Based NLP Tools for Domain-Specific Text in Foreign Languages

A structured framework combining multi-system machine translation evaluation and active learning for domain-specific text classification for domain-specific foreign-language corpora is provided and empirically validate.

Naif Alatrush, Luay Abdeljaber, Javier Osorio et al. · 0 citations
Open access Aug 2026

Low-Resource Machine Translation of Yoruba Text to Nigerian Pidgin Using NLLB-200 and Parameter-Efficient Fine-Tuning

Machine translation for low-resource and structurally divergent languages remains a significant challenge, particularly when mapping highly tonal languages to contact languages with fluid orthographies. This study presents the development of a translation system for Yoruba to Nigerian Pidgin, which addresses a critical gap in African natural language processing. Using a custom-curated parallel corpus of 1,846 pairs, the NLLB-200 (600M) foundational model was adapted for this specific language pair. To reduce the computational cost of full-model fine-tuning, Low-Rank Adaptation (LoRA), a Parameter-Efficient Fine-Tuning (PEFT) technique, was applied to the transformer's attention layers. Applying systematic hyperparameter ablation, the model was evaluated across various LoRA ranks and beam search decoding sizes. The results from the evaluation demonstrate that the optimal configuration (LoRA rank 16, Beam Size 4) significantly outperformed the zero-shot baseline (paired bootstrap resampling, p < 0.001), improving the BLEU (Bilingual Evaluation Understudy) score from 0.74 to 30.97 and the character-level score from 9.59 to 50.63. Further qualitative evaluations confirmed the model's ability to generate semantically adequate and colloquially natural translations despite the lack of a standardized dictionary for Nigerian Pidgin. These findings demonstrate the efficacy of parameter-efficient adaptation strategies in democratizing translation technologies for underrepresented, non-standardized linguistic environments.

Oyebola Akande, G. Opateye · 0 citations
Open access Jul 2026

Transliteration for Low-Resource Translation in the Age of Large Language Models

The findings show that manual transliteration consistently yields the best translation performance, while noisy automatic romanization reduces these gains, and that LLM-based translation can be competitive with, and in some settings outperform, fine-tuned NMT systems, although this advantage comes with lower interpretability.

A. Mansurova, Meruert Bekmukhamedova, Bekarys Baibolat et al. · 0 citations
Jul 2026

Systematic Analysis of Large Language Models and Transformer-Based Machine Translation for English-Tamil and Tamil-English Across Diverse Datasets

The findings show that the quality of the datasets and their alignment with the domain will greatly affect the performance of the model, that attention-based mechanisms can aid in explain ability, and that few-shot large language models can still produce structurally coherent translations of Tamil.

S. Sriharshaa, Sangeetha Sivanesan, S. JayaNirmala · 0 citations
Open access Jul 2026

Automated Multilingual Translator Using Neural Translation

The results indicate that a moderately sized, shared self-attention architecture can deliver production-quality multilin-gual translation within the resource constraints of an academic de-ployment, while surfacing clear directions – low-resource language coverage, domain adaptation, and speech-based extension – for con-tinued development.

Darshan Gowda D H and Dr. Kruti R · 0 citations
Open access Sep 2026

SSPA: Enhancing Pseudo-Corpus Quality on Tibetan Machine Translation via Semantic-Syntax Prealignment

Tibetan-to-English machine translation (MT) models frequently falter under extreme domain data scarcity, often producing translations that violate the distinctive agglutinative rules of Tibetan and suffer from domain-specific stylistic mismatches. To overcome these limitations, we propose Semantic-Syntax Prealignment (SSPA), an innovative corpus generation framework. SSPA constructs high-quality pseudo-parallel pairs by explicitly minimizing the deviation between the syntactic-semantic profiles of generated samples and professional reference texts. Specifically, source-target structural representations are standardized through length-unified truncation and terminology normalization, followed by a dual-domain alignment process that maximizes syntactic cosine similarity under rigorous structural constraints. We further augment these aligned frames via a cross-length dynamic filling mechanism, which is integrated with an Expectation-over-Transformation (EOT)-based style regularization mechanism specifically adapted for stylistic perturbations, to simulate authentic linguistic variations. Extensive evaluations on our newly constructed Tibetan Medicine-Tibetan English (TM-TE) dataset demonstrate that SSPA significantly outperforms existing competitive baselines. Notably, SSPA achieves a BLEU-4 score of 36.2 and improves long-sentence BLEU-4 by 16.8 points, with a parser-verified grammatical compliance rate of 96.2%. The framework exhibits remarkable cross-domain adaptability and stylistic consistency, offering a robust, versatile solution for low-resource Tibetan professional domain MT.

Yi-Dong Sun, Dong-Xu Liu, Jia-Lei Zhang et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.