Skip to content
Open access

The Shortcomings of Natural Language Processing (NLP) Models and Their Applications in Achieving Accuracy in Translating from Arabic to English

Sep 2026 · Digital Technologies Research and Applications · 0 citations · 28 references

TL;DR

This study highlights the importance of integrating in-depth linguistic analysis, contextual semantic modelling, and cultural awareness into natural language processing-based translation systems by combining traditional linguistic insights with computational methods.

Abstract

Translation is defined as the process of transferring meaning from one language to another. It is an extremely difficult and complex process because it involves not only transferring words but also ideas, culture, linguistic customs, and meanings derived from syntactic elements, their arrangement, word structure, and derivation. This is especially true in languages with complex structures, such as Arabic, which is characterized by its multiple linguistic contexts. These contexts have not been adequately addressed by NLP (Natural Language Processing) applications in machine translation models due to the lack of diverse contexts where metaphor, figurative language, grammatical inflections, morphological patterns and their connotations, and sentence structure all play pivotal roles in determining meaning. Furthermore, spoken language, with its inherent phonetic and expressive characteristics, conveys the text into broader semantic spaces. These spaces are influenced by the effect of intonation on specific syllables, the speaker's psychological state, the listener's mood, and accompanying body language, which transforms meaning into other subtle details. All of this, and more, is absent from machine translation, no matter how hard its creators try to imbue it with human emotions and feelings. This study highlights the importance of integrating in-depth linguistic analysis, contextual semantic modelling, and cultural awareness into natural language processing-based translation systems. By combining traditional linguistic insights with computational methods, the research offers a framework that can contribute to improving the accuracy of machine translation from Arabic to English.

Read PDF

Similar papers

Open access Jul 2026

A UNIFIED LINGUISTIC AWARE PRE-PARSING FRAMEWORK FOR ENRICHING ENGLISH TO INDIAN MACHINE TRANSLATION

Machine Translation has become one of the major application areas of Artificial Intelligence (AI) and Natural Language Processing (NLP), especially in multilingual countries like India. Although recent Neural Machine Translation systems have shown good performance for several language pairs, translation quality is still inconsistent for many Indian languages because of linguistic and structural differences between English and Indian language families. Most Indian languages are morphologically rich and contain flexible word order, complex agreement patterns, compound constructions, and context-dependent grammatical forms. Because of this, direct translation from English often produces structurally incorrect or semantically weak output. In many existing systems, the source sentence is passed to the translation model without sufficient linguistic analysis. As a result, ambiguity present in the source text propagates further during translation. This work focuses on the importance of linguistic enrichment before the translation stage. The proposed framework, named Unified Linguistic-Aware Pre-Parsing Framework, introduces a coordinated pre-processing layer for English-to-Indian Machine Translation (MT). A key contribution of this research is the development of a novel linguistically enriched intermediate representation that extends beyond conventional text normalization. By transforming noisy input text into linguistically enriched translation-ready representation, the proposed approach facilitates effective knowledge transfer to machine translation models, leading to improve contextual adequacy, linguistic fidelity, and overall translation performance. The framework combines multiple linguistic processing stages including POS tagging, NE detection, clause boundary analysis, contextual token handling, syntactic structure preparation, and morphology-related processing. Instead of executing these modules independently, the proposed system allows interaction between lexical, syntactic, and morphological information during analysis. This helps reduce structural ambiguity and improves sentence-level interpretation before translation begins. The need for such a framework becomes more relevant in the context of Indian languages where morphology and grammatical relations carry significant semantic information. This framework is especially relevant for Indian languages, where semantic information is often encoded through morphological variations and grammatical dependencies. The proposed framework can be effectively integrated with both conventional machine translation architectures and modern large language models. The overall study highlights how classical linguistic analysis can still play an important role in improving multilingual AI systems for Indian languages.

Prashant Chaudhary, Pavan Kurariya, Jahnavi Bodhankar et al. · 0 citations
Preprint Aug 2026

An Investigation of Translationese in the Generations of Multilingual Large Language Models

Text which has been translated from another language tends to carry with it evidence of translation$\unicode{x2014}$hence, it is often referred to as $\textit{translationese}$. Multilingual large language models (MLLMs) generate text in a variety of languages. However, it is still unclear if MLLMs'generations resemble internal translation (from English or, potentially, other languages) and, thus, result in translationese. Here, we ask the following research questions: (1) Does text generated by MLLMs resemble translationese? (2) How does translationese produced by MLLMs differ from translationese produced through direct translation? We leverage established indicators of translated text to evaluate text generated by state-of-the-art MLLMs in five languages, comparing to both non-translated and human-written baselines in order to isolate translationese from other kinds of interference. Through the use of high-accuracy classification models, analyses of variance on individual linguistic features, and the collection of human annotations in a subset of two languages (German and Spanish), we assess the translationese content of MLLM generations and examine the key features that distinguish MLLM-generated text from typical translation-related interference.

Maria R. Valentini, Téa Wright, Julisa Granados et al. · 0 citations
Open access Sep 2026

Bridging the linguistic divide: recent developments in machine translation for Indian languages

This paper analyses various recent state-of-the-art variants of large language models (LLMs) and neural machine translation (NMT) for Indian languages in comparison to statistical machine translation (SMT) and tackles key questions, such as idiomatic expressions, morphologically complex grammar or the scarceness of parallel corpora.

Jayanand A. Kamble, S. Jadhav, V. J. Kadam · 0 citations
Open access Aug 2026

Linguistic analysis of structural and textual ambiguity in english

The findings provide insights into the mechanisms employed by large language models in processing ambiguous language and demonstrate similarities and differences in their contextual interpretation of structurally ambiguous constructions.

Khanbutayeva Leyla Musa · 0 citations
Open access Aug 2026

Using DeepSeek as a Russian-Chinese Translation Tool

The subject of the research is the semantic, grammatical, and pragmatic characteristics of translations generated by DeepSeek in comparison with DeepL and Google Translate. The object is machine translation in the Russian–Chinese language pair using generative language models. The relevance is determined by the contradiction between the expanding use of large language models in translation and the insufficiently defined boundaries of their effectiveness with typologically and culturally distant languages. The aim is to determine the boundaries of DeepSeek's effective application in both translation directions. The objectives include characterizing the model's technological features, comparing cognitive mechanisms of language processing by humans and artificial intelligence, and empirically testing translation quality against DeepL and Google Translate according to semantic accuracy, grammatical correctness, and pragmatic adequacy. The study employed comparative analysis, cognitive modeling, and interpretive analysis on a corpus of 30 phraseological and culturally marked units. The scientific novelty lies in the systematization of knowledge about generative neural networks in translation theory and in the comparative assessment of three systems on a unified corpus according to three complementary criteria. The author's contribution consists in refining the understanding of similarities and differences between human and machine cognitive mechanisms. The main findings are as follows: DeepSeek outperforms DeepL and Google Translate in conveying idioms and cultural realia, however its functional adaptation may lead to semantic shifts, necessitating professional post-editing in terminologically dense texts. The most justified application is producing draft translations and finding contextual equivalents, while final verification should remain with the human translator.

Ilia Alekseevich Konstantinov · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.