Jul 2026· International Journal of Data Science and Analysis· Vol 22· 0 citations· 54 references
Computer Science
TL;DR
The results show that LLMs and Google Translate consistently outperform specialized MT systems in terms of fluency, meaning preservation, and lexical-thematic alignment.
Abstract
Automatic poetry translation remains a challenging task, as it requires not only semantic accuracy but also the preservation of stylistic and emotional elements. This study investigates the effectiveness of machine translation (MT) systems and large language models (LLMs) in poetry translation. Traditional automatic evaluation metrics often fail to capture the literary quality of such translations; therefore, we propose a three-phase evaluation framework that integrates complementary perspectives: (i) automatic metrics (BLEU, METEOR, and BERTScore) to assess lexical and semantic fidelity, (ii) topic modeling with BERTopic to perform an exploratory analysis of lexical-thematic alignment across translations, and (iii) expert human evaluation to examine poetic structure, style, fluency, and meaning preservation. This framework was applied to compare specialized MT systems (mBART, MarianMT, OpenNMT with RNN, and Google Translate) with LLMs such as ChatGPT and Maritaca AI across six language pairs involving English, French, and Portuguese (English–French, English–Portuguese, French–English, French–Portuguese, Portuguese–English, and Portuguese–French). Additionally, we investigate the impact of fine-tuning strategies using corpora of poems and song lyrics. The results show that LLMs and Google Translate consistently outperform specialized MT systems in terms of fluency, meaning preservation, and lexical-thematic alignment. However, human evaluation reveals that all systems struggle to replicate the poetic structure and stylistic nuances of the originals. The fine-tuning process did not produce improvements for all models; mBART showed notable gains, while the other models did not benefit significantly from domain adaptation.
The paper argues that NLP should operate as an interpretive assistant rather than an autonomous literary translator in translating Iraqi poetry into English, and proposes a culturally aware, human-in-the-loop framework for supporting literary translation.
Whaj Mneer Esmail· Iraqi Literary and Cultural...· 0 citations
This article examines the quality of AI-generated poetry translation through a comparative study of Wilfred Owen’s Dulce et Decorum Est translated from English into Italian. The analysis compares Sergio Rufini’s published human translation with two AI-generated versions produced by ChatGPT and Google Translate. The study combines qualitative stanza-by-stanza close reading with a reduced Multidimensional Quality Metrics framework in order to assess both identifiable translation errors and broader losses of poetic, intercultural, and rhetorical force. The findings show that both AI-generated translations remain below the quality thresholds established in the MQM scorecard. ChatGPT produces a fluent and relatively coherent Italian version, but its output shows significant problems of undertranslation and poetic regularisation. Google Translate obtains a higher calibrated MQM score in this dataset, but its translation remains strongly oriented toward sentence-level transfer and does not consistently preserve the poem’s cumulative rhetorical structure. The article argues that AI systems can produce linguistically plausible literary translations, but that poetic adequacy requires interpretive prioritisation, intercultural judgement, and aesthetic decision-making. The study contributes to research on AI-assisted literary translation by distinguishing surface fluency from poetic and intercultural adequacy.
Large language models (LLMs) have rapidly expanded the horizon of AI-assisted literary translation, yet the mechanisms of how prompts construct translation purpose are theoretically underexplored. This study, based on Vermeer’s Skopos Theory, conceptualizes prompts as a “digital translation brief” and examines the systematic regulation of four function-oriented prompts: baseline (P0), reader-oriented (P1), form-oriented (P2), and culture-oriented (P3) in DeepSeek’s translation strategies for English-to-Indonesian poetry translation, with Emily Dickinson’s Hope is the Thing with Feathers as the source text. The study finds, through comparative close reading across the four prompt conditions, as well as against two published human translations, that each prompt orientation produces distinct and observable shifts in diction, imagery construction and formal expression: P1 sacrifices collocational stability for literary register and emotional intensity at the expense of; P2 replicates structure but risks over-compliance which undermines target-reader adequacy; P3 allows discourse-level metaphorical coordination and selective dependence on the source-text. Human translators’ choices emerge from cultural memory, aesthetic intentionality, and poetic agency. DeepSeek’s functional adaptability is largely a matter of probabilistic generation and prompt compliance. The findings re-conceptualize prompts not only as technical input instructions, but as purposive regulators of translation, offering a translation-theoretic framework for analysing LLM behaviour and practical guidance for designing Skopos-informed prompts for AI-assisted literary translation.
Haijin Li· Journal of Translation and L...· 0 citations
This study analyzes varied Chinese renderings of the Buddhist core concept CITTA via quantitative modeling to expose cross-cultural translation discrepancies. It constructs a BEiT-RoBERTa fusion model combining BEiT visual feature extraction and RoBERTa linguistic encoding. Based on a self-built CITTA Chinese translation text dataset containing ancient and modern translations, experiments were conducted from four dimensions: vocabulary recognition, model comparison, semantic coherence, and multi-context adaptation. Performance tests were conducted using indicators such as character error rate (CER), syllable recognition accuracy (SRA), line recognition accuracy (LRA), and manual evaluation. Results show the syllable-based model performs best: CER=0.167, SRA=0.850, LRA=0.791, surpassing CRNN and SRN. The overall accuracy of sentence semantic coherence judgment is 92.68%, and the translation effect in different contexts is between professional translators and non professional enthusiasts, with stable accuracy and fluency. Research has confirmed that the BEiT-RoBERTa model quantifies CITTA translation variations but poorly captures philosophical depth and cultural subtleties.
Xiaoming Chen· WSEAS Transactions on Comput...· 0 citations
This paper examines the features and difficulties of machine translation of William Shakespeare’s poetic texts, in particular, cases of distortion or violation of lexical units and stable phrases. The analysis is based on a comparison of the original text with the results of a translation performed using modern machine translation systems. Special attention is paid to typical errors related to the ambiguity of words, violation of syntactic connections, ignoring the cultural and contextual meaning of metaphors, as well as the destruction of poetic structure. The work highlights the need to take into account the stylistic and semantic features of poetic discourse in automated translation and suggests ways to improve machine translation algorithms to more accurately convey the artistic and emotional nuances of poetry.
D.T. Teshebaeva· Vestnik of the Kyrgyz-Russia...· 0 citations
As a highly expressive rhetorical device in classical Chinese poetry, reduplicated words possess unique aesthetic value in semantic connotation, phonological rhythm and morphological presentation. However, due to the fundamental differences between Chinese and English linguistic systems, the English translation of Chinese reduplicated words has long been a difficult task in translation practice. This paper looks at two English translations of the famous reduplicative opening in Li Qingzhao’s Slow, Slow Tune. Using eco-translatology as its theoretical lens, it compares the translators’ choices, examining how each navigated the constraints of language, culture, and communicative effect. What emerges is a picture of translation as a dynamic balancing act—a process in which the translator constantly adjusts, compromises, and seeks the best fit within a complex web of forces. The findings offer a systematic way of thinking about how reduplicated words in classical poetry might be translated and evaluated more thoughtfully.
Xuri Shen· English Language Teaching an...· 0 citations