Skip to content
Review Open access

Evaluating Terminological Consistency in AI-Generated English–Arabic Political Translation

Aug 2026 · (Faculty of Arts Journal) مجلة كلية الآداب - جامعة مصراتة · 0 citations · 18 references

TL;DR

The study concludes that terminological consistency should be considered as a separate dimension of translation quality and emphasizes the importance of terminology control and human post-editing in AI-assisted political translation.

Abstract

Machine translation has become more fluent and contextually accurate with recent advances in artificial intelligence. However, terminological consistency has been underexplored, particularly in political and electoral discourse where lexical repetition and conceptual precision are critical for cohesion and clarity. This study investigates terminological consistency in AI-generated English–Arabic political translations produced by ChatGPT and Google Gemini Advanced. The translations were generated and analyzed between January and June 2026 using the systems’ default settings to ensure comparability and avoid potential variations resulting from user-configured parameters. The study employs a mixed-methods corpus-based approach. The study analyzes 30 political and electoral texts with 60 recurring key terms. Quantitative analysis measures the degree of consistency in the form of stability percentages, and qualitative analysis studies lexical variation and its effect on discourse cohesion and clarity. The adequacy and consistency of the translation were checked against a reference translation based on the United Nations Development Programme (UNDP) Arabic Lexicon of Electoral Terminology. The results indicate that ChatGPT achieved higher terminological consistency than Google Gemini. ChatGPT’s lexical equivalents for repeated political and electoral terms were more stable than its Gemini counterpart, which showed more lexical variation, especially in context-sensitive terms such as campaign, electoral law, and judicial review. The study concludes that terminological consistency should be considered as a separate dimension of translation quality and emphasizes the importance of terminology control and human post-editing in AI-assisted political translation

Read PDF

Similar papers

Open access Aug 2026

Zoonym-Based Phraseology in English–Azerbaijani AI-Assisted Translation: Semantic, Affective, and Pragmatic Equivalence

The study demonstrates the value of context-sensitive, linguoculturally informed evaluation of AI-assisted phraseological translation by illustrating successful preservation of conventional target-language equivalents in some items and literal rendering, metaphorical-image mismatch, reduced emotional expressiveness, pragmatic weakening, or loss of cultural symbolism in others.

Aysel V. Safarova · 0 citations
Open access Jul 2026

Evaluating Artificial Intelligence Tools In Identifying Translation Techniques In Literary Texts

This study concludes that while all AI tools can transfer denotative meaning, generative models (LLMs) are superior in applying high-level translation techniques necessary for maintaining emotional nuance, discourse cohesion, and literary appeal.

Ika Oktaria, Abitya Sakti Fathan Narotama · 0 citations
Open access Jul 2026

Assessing English-Arabic translation of verb phrase ellipsis: A comparative study of Google Translate and ChatGPT-4o

Assessment of how GT and GPT-4o translate English VPE into Arabic focuses on the accuracy of ellipsis reconstruction and the translation strategies employed, highlighting the importance of better-quality training data for NMT and LLM tools for both discourse-level processing and context-dependent data.

Eassa Ali, Abbas Brashi, Dana Awad et al. · 0 citations
Open access Jul 2026

Large Language Models for Japanese–Croatian Translation: Human Evaluation and Macroeconomic Implications

Large Language Models (LLMs) are increasingly used for translation, yet their value depends on preserving meaning rather than producing fluent output. This study evaluates seven LLMs on Japanese–Croatian translation, a low-resource, typologically distant language pair. Using rubric-based human evaluation of adequacy, fluency, terminology, and register, we compare model performance. Results show a stable ranking: qwen3 performs best, followed by phi4 and gemma3, while qwen2 performs worst. Performance differences reflect structural reconstruction, particularly argument recovery, aspectual mapping, lexical precision, and register. Qualitative analysis also reveals limited differentiation within the South Slavic continuum and pragmatic inconsistencies. Although productivity effects were not measured, improved translation adequacy may reduce post-editing and verification effort.

Ratomir Karlović, Mieta Bobanović Dasko, Irena Srdanović · 0 citations
Open access Aug 2026

Using DeepSeek as a Russian-Chinese Translation Tool

The subject of the research is the semantic, grammatical, and pragmatic characteristics of translations generated by DeepSeek in comparison with DeepL and Google Translate. The object is machine translation in the Russian–Chinese language pair using generative language models. The relevance is determined by the contradiction between the expanding use of large language models in translation and the insufficiently defined boundaries of their effectiveness with typologically and culturally distant languages. The aim is to determine the boundaries of DeepSeek's effective application in both translation directions. The objectives include characterizing the model's technological features, comparing cognitive mechanisms of language processing by humans and artificial intelligence, and empirically testing translation quality against DeepL and Google Translate according to semantic accuracy, grammatical correctness, and pragmatic adequacy. The study employed comparative analysis, cognitive modeling, and interpretive analysis on a corpus of 30 phraseological and culturally marked units. The scientific novelty lies in the systematization of knowledge about generative neural networks in translation theory and in the comparative assessment of three systems on a unified corpus according to three complementary criteria. The author's contribution consists in refining the understanding of similarities and differences between human and machine cognitive mechanisms. The main findings are as follows: DeepSeek outperforms DeepL and Google Translate in conveying idioms and cultural realia, however its functional adaptation may lead to semantic shifts, necessitating professional post-editing in terminologically dense texts. The most justified application is producing draft translations and finding contextual equivalents, while final verification should remain with the human translator.

Ilia Alekseevich Konstantinov · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.