Skip to content
Open access

Using English-Based NLP Tools for Domain-Specific Text in Foreign Languages

2026 · IEEE Access · Vol 14, pp. 111122-111139 · 0 citations · 101 references
Computer Science

TL;DR

A structured framework combining multi-system machine translation evaluation and active learning for domain-specific text classification for domain-specific foreign-language corpora is provided and empirically validate.

Abstract

Social scientists often machine-translate foreign-language texts into English and apply English-based natural language processing tools without systematically evaluating translation quality or annotation efficiency. To address this problem, this study provides evidence-based guidance for researchers applying English-centric natural language processing to domain-specific foreign-language corpora. We provide and empirically validate a structured framework combining multi-system machine translation evaluation and active learning for domain-specific text classification. Using 11,493 parallel Spanish and Arabic sentences aligned to English, we compare four machine translation systems (Google Translate, Deep, DeepL, OPUS) using SacreBLEU, METEOR, COMET, and BERTScore quality scores. Across languages and metrics, machine translation systems yield statistically comparable performance. We then evaluate eight active learning strategies using ConfliBERT for political conflict classification under a 20% annotation budget, corresponding to 1,155 samples from the training split. Binary classification exceeds F1 = 0.90, while QuadClass multi-class performance peaks around F $1~\approx ~0.75$ . The Ensemble Intersection strategy achieves the highest performance in 53% of tasks and often matches or surpasses full-dataset results using only a fraction of labeled data. These results provide a practical workflow for researchers using English-based natural language processing tools on foreign-language, domain-specific corpora.

Read PDF

Similar papers

Preprint Jul 2026

Translation as a Computationally Efficient Bridge: Feasibility of English BERT for Low-Resource Languages

Evaluating the feasibility of translation-based fine-tuning across six NLP tasks demonstrates that translation-based fine-tuning offers a scalable, resource-efficient, and empirically validated path for extending NLP to low-resource languages while advancing linguistic inclusivity and sustainability in artificial intelligence.

H. Muizelaar, Giulia Rivetti, Marco Spruit et al. · 0 citations
Review Open access Aug 2026

Multi-Metric Evaluation of Translation-Based Cross-Lingual Sentiment Consistency Using Large Language Models and Neural Machine Translation

In today’s globalized and digitally connected world, individuals increasingly share emotions, opinions, and experiences across multiple languages, making accurate translation essential for cross-lingual sentiment analysis. Although machine translation (MT) is widely used in multilingual applications, the relationships among translation quality, semantic similarity, and sentiment consistency remain insufficiently understood. This study investigates the performance of six LLM-based systems (GPT-4o-mini, Gemini 2.5 Flash-Lite, Qwen 2.5, Llama 3.1, Mistral 7B, and NiuTrans LMT) and four NMT-based systems (Google Translate, Microsoft Translator, NLLB-200, and LibreTranslate-v1.5) in maintaining classifier-mediated sentiment consistency across twelve translation directions involving English, Spanish, French, and Chinese. Experiments were conducted on the Multilingual Amazon Reviews Corpus (MARC), comprising 84,000 randomly sampled user reviews. A multidimensional evaluation framework was used, combining sentiment-consistency metrics (accuracy, weighted F1, MCC, and SSR), translation-quality estimation (COMET-QE), and semantic-similarity assessment (LaBSE). Statistical significance was examined using the Friedman and Nemenyi post hoc tests. The results show that GPT-4o-mini, Gemini 2.5 Flash-Lite, and Google Translate consistently ranked among the strongest systems across multiple evaluation dimensions. Performance differences were particularly pronounced in translation directions involving Chinese, highlighting the influence of language-specific structural characteristics. Furthermore, semantic similarity and translation quality exhibited only moderate relationships with sentiment consistency, indicating that high semantic similarity does not necessarily guarantee strong sentiment consistency. Overall, the findings demonstrate the importance of multidimensional and statistically grounded evaluation frameworks for assessing cross-lingual sentiment consistency and provide practical insights into the strengths and limitations of contemporary MT systems.

E. Cetin, Çağrı Şahin · 0 citations
Jul 2026

Systematic Analysis of Large Language Models and Transformer-Based Machine Translation for English-Tamil and Tamil-English Across Diverse Datasets

The findings show that the quality of the datasets and their alignment with the domain will greatly affect the performance of the model, that attention-based mechanisms can aid in explain ability, and that few-shot large language models can still produce structurally coherent translations of Tamil.

S. Sriharshaa, Sangeetha Sivanesan, S. JayaNirmala · 0 citations
Open access Jul 2026

A UNIFIED LINGUISTIC AWARE PRE-PARSING FRAMEWORK FOR ENRICHING ENGLISH TO INDIAN MACHINE TRANSLATION

Machine Translation has become one of the major application areas of Artificial Intelligence (AI) and Natural Language Processing (NLP), especially in multilingual countries like India. Although recent Neural Machine Translation systems have shown good performance for several language pairs, translation quality is still inconsistent for many Indian languages because of linguistic and structural differences between English and Indian language families. Most Indian languages are morphologically rich and contain flexible word order, complex agreement patterns, compound constructions, and context-dependent grammatical forms. Because of this, direct translation from English often produces structurally incorrect or semantically weak output. In many existing systems, the source sentence is passed to the translation model without sufficient linguistic analysis. As a result, ambiguity present in the source text propagates further during translation. This work focuses on the importance of linguistic enrichment before the translation stage. The proposed framework, named Unified Linguistic-Aware Pre-Parsing Framework, introduces a coordinated pre-processing layer for English-to-Indian Machine Translation (MT). A key contribution of this research is the development of a novel linguistically enriched intermediate representation that extends beyond conventional text normalization. By transforming noisy input text into linguistically enriched translation-ready representation, the proposed approach facilitates effective knowledge transfer to machine translation models, leading to improve contextual adequacy, linguistic fidelity, and overall translation performance. The framework combines multiple linguistic processing stages including POS tagging, NE detection, clause boundary analysis, contextual token handling, syntactic structure preparation, and morphology-related processing. Instead of executing these modules independently, the proposed system allows interaction between lexical, syntactic, and morphological information during analysis. This helps reduce structural ambiguity and improves sentence-level interpretation before translation begins. The need for such a framework becomes more relevant in the context of Indian languages where morphology and grammatical relations carry significant semantic information. This framework is especially relevant for Indian languages, where semantic information is often encoded through morphological variations and grammatical dependencies. The proposed framework can be effectively integrated with both conventional machine translation architectures and modern large language models. The overall study highlights how classical linguistic analysis can still play an important role in improving multilingual AI systems for Indian languages.

Prashant Chaudhary, Pavan Kurariya, Jahnavi Bodhankar et al. · 0 citations
Open access Jul 2026

Automated Multilingual Translator Using Neural Translation

The results indicate that a moderately sized, shared self-attention architecture can deliver production-quality multilin-gual translation within the resource constraints of an academic de-ployment, while surfacing clear directions – low-resource language coverage, domain adaptation, and speech-based extension – for con-tinued development.

Darshan Gowda D H and Dr. Kruti R · 0 citations
Open access Aug 2026

A Data-Efficient Multilingual Neural Machine Translation Model for Low-Resource Indic Languages

The effectiveness of multilingual transfer learning in low-resource settings is demonstrated by the fine-tuned Multilingual Bidirectional and Auto-Regressive Transformer-50 model, significantly outperforming the pretrained baseline.

G. Harshitha, Vasudeva, Nisha P. Poojary et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.