Skip to content
Preprint

Translation as a Computationally Efficient Bridge: Feasibility of English BERT for Low-Resource Languages

Jul 2026 · 0 citations · 82 references
Computer Science

TL;DR

Evaluating the feasibility of translation-based fine-tuning across six NLP tasks demonstrates that translation-based fine-tuning offers a scalable, resource-efficient, and empirically validated path for extending NLP to low-resource languages while advancing linguistic inclusivity and sustainability in artificial intelligence.

Abstract

BERT models have revolutionised Natural Language Processing (NLP) through their ability to process unstructured text across diverse domains. However, developing high-quality BERT models for non-English languages remains challenging due to limited annotated data and high computational demands. Translating non-English data into English and fine-tuning existing English BERT models offers a resource-efficient alternative, yet few studies have structurally compared translation-based fine-tuning with native-language BERT performance across tasks and languages. This study provides such a comparison, evaluating the feasibility of translation-based fine-tuning across six NLP tasks: Sentiment Analysis, Hate Speech Detection, Question Answering, Named Entity Recognition, Part-of-Speech Tagging, and Natural Language Inference, using datasets translated from Bulgarian, Chinese, Dutch, Italian, and Russian. Across all settings, the translation-based approach was comparable or superior in 53.3 percent of cases. Gains were most frequent in Question Answering, Part-of-Speech Tagging, and Natural Language Inference, while performance declines were common in Named Entity Recognition and Hate Speech Detection. The results show that translation-based fine-tuning is most effective for tasks relying on syntactic or structural patterns and for languages typologically close to English, such as Dutch, but less effective for token-level or culturally nuanced tasks, particularly in Chinese. Overall, this study demonstrates that translation-based fine-tuning offers a scalable, resource-efficient, and empirically validated path for extending NLP to low-resource languages while advancing linguistic inclusivity and sustainability in artificial intelligence.

View source

Similar papers

Open access 2026

Using English-Based NLP Tools for Domain-Specific Text in Foreign Languages

A structured framework combining multi-system machine translation evaluation and active learning for domain-specific text classification for domain-specific foreign-language corpora is provided and empirically validate.

Naif Alatrush, Luay Abdeljaber, Javier Osorio et al. · 0 citations
Open access Jul 2026

Transliteration for Low-Resource Translation in the Age of Large Language Models

The findings show that manual transliteration consistently yields the best translation performance, while noisy automatic romanization reduces these gains, and that LLM-based translation can be competitive with, and in some settings outperform, fine-tuned NMT systems, although this advantage comes with lower interpretability.

A. Mansurova, Meruert Bekmukhamedova, Bekarys Baibolat et al. · 0 citations
#natural language process... Preprint Jul 2026

A Factorial Study of Synthetic Data Generation for Low-Resource Machine Translation using Grammar Books

A pipeline that uses large language models to extract grammatical rules, example sentences, and lexicons from grammar books and generate synthetic parallel corpora for fine-tuning-rather than feeding grammar content into prompts at inference time, as in prior work is introduced.

V. Ravikumar, Sina Ahmadi, L. Jäger et al. · 0 citations
Open access Sep 2026

Bridging the linguistic divide: recent developments in machine translation for Indian languages

This paper analyses various recent state-of-the-art variants of large language models (LLMs) and neural machine translation (NMT) for Indian languages in comparison to statistical machine translation (SMT) and tackles key questions, such as idiomatic expressions, morphologically complex grammar or the scarceness of parallel corpora.

Jayanand A. Kamble, S. Jadhav, V. J. Kadam · 0 citations
Open access Jul 2026

Croatian Language in the Transition from Neural Machine Translation to Large Language Models

The evaluation of various NMT and LLM architectures specifically for Croatian from/to English and Spanish demonstrates that open–source models can achieve, and occasionally surpass, the quality of Google Translate, a widely used commercial NMT system.

Antoni Oliver, Sergi Álvarez–Vidal · 1 citation

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.