Intercultural Competence and Translation Quality: Assessing AI and Human Literary Translations
Abstract
This article examines the quality of AI-generated poetry translation through a comparative study of Wilfred Owen’s Dulce et Decorum Est translated from English into Italian. The analysis compares Sergio Rufini’s published human translation with two AI-generated versions produced by ChatGPT and Google Translate. The study combines qualitative stanza-by-stanza close reading with a reduced Multidimensional Quality Metrics framework in order to assess both identifiable translation errors and broader losses of poetic, intercultural, and rhetorical force. The findings show that both AI-generated translations remain below the quality thresholds established in the MQM scorecard. ChatGPT produces a fluent and relatively coherent Italian version, but its output shows significant problems of undertranslation and poetic regularisation. Google Translate obtains a higher calibrated MQM score in this dataset, but its translation remains strongly oriented toward sentence-level transfer and does not consistently preserve the poem’s cumulative rhetorical structure. The article argues that AI systems can produce linguistically plausible literary translations, but that poetic adequacy requires interpretive prioritisation, intercultural judgement, and aesthetic decision-making. The study contributes to research on AI-assisted literary translation by distinguishing surface fluency from poetic and intercultural adequacy.