Aug 2026· The social science· 0 citations· 16 references
TL;DR
It is suggested that LLM agents can reach a strong MT baseline and may offer pragmatic and audience-oriented affordances, while their value depends on the benchmark system, the target discourse function, and the evaluation dimension.
Abstract
This study evaluates how web-based machine translation (MT) systems and large language model (LLM) translation agents perform in the English translation of Chinese city publicity texts. City publicity translation is a high-stakes form of institutional intercultural communication: it must be factually accurate, culturally legible, pragmatically appropriate, accessible to international readers, and capable of representing a city image without exaggeration or distortion. A corpus of 90 official Chinese source segments from Qingdao, Xi'an, and Hangzhou was translated under four conditions: DeepL web MT, Google Translate web MT, a GPT-5.5 translation agent, and a DeepSeek V4 Pro translation agent, yielding 360 English translations. Two trained coders independently evaluated all translations on five 1-5 dimensions: accuracy, cultural adequacy, pragmatic appropriateness, audience accessibility, and city image representation. Formal coding showed acceptable to strong reliability: Cohen's kappa for primary issue coding was 0.777, and quadratic weighted kappa values for the five rating dimensions ranged from 0.786 to 0.847. The strongest composite score was observed for DeepL web MT (M=4.683), followed by GPT-5.5 (M=4.599), DeepSeek V4 Pro (M=4.553), and Google Translate (M=4.261). Paired permutation tests showed that DeepL, GPT-5.5, and DeepSeek V4 Pro all significantly outperformed Google Translate on composite quality, while the differences between DeepL and the two LLM agents were not statistically significant. The findings therefore do not support a simple claim that LLM agents uniformly surpass MT. Instead, they suggest that LLM agents can reach a strong MT baseline and may offer pragmatic and audience-oriented affordances, while their value depends on the benchmark system, the target discourse function, and the evaluation dimension.
Assessing the quality of scientific literature translation remains challenging because of strong subjectivity, dense domain-specific terminology, and the limited availability of standardized reference translations. These issues are particularly relevant for the international dissemination of research in advanced electromagnetic engineering, where precise multilingual communication supports the reliable exchange of knowledge on electromagnetic waves, antennas, and propagation technologies. This paper proposes a Contrastive Learning-based Chinese-English Scientific Translation Quality Evaluation model (C-TQE). By constructing multi-level positive and negative sample pairs, the model learns the relative ordinal relationships of translation quality within a shared representation space. A dual-encoder architecture encodes source sentences and candidate translations through a shared pre-trained language model, while a contrastive loss function draws high-quality translations closer to the source representation and separates low-quality ones. To address the characteristics of scientific texts, a term-aware negative sampling strategy exploits domain dictionaries and syntactic structures to generate semantically similar but terminologically incorrect examples. Experiments on 11, 238 human-annotated instances from the WMT20–22 Chinese-English scientific translation tasks show that C-TQE achieves a Kendall’s tau correlation coefficient of 0.564 with human judgments, outperforming COMET (0.512) and BLEURT (0.497). Ablation studies confirm the effectiveness of term-aware negative sampling and the contrastive learning objective, while diagnostic analysis demonstrates high consistency in evaluating terminological accuracy and syntactic structures. The proposed framework provides an effective solution for large-scale scientific translation quality assessment and facilitates the accurate international communication of multidisciplinary engineering research, including electromagnetic and antenna-related studies.
In today’s globalized and digitally connected world, individuals increasingly share emotions, opinions, and experiences across multiple languages, making accurate translation essential for cross-lingual sentiment analysis. Although machine translation (MT) is widely used in multilingual applications, the relationships among translation quality, semantic similarity, and sentiment consistency remain insufficiently understood. This study investigates the performance of six LLM-based systems (GPT-4o-mini, Gemini 2.5 Flash-Lite, Qwen 2.5, Llama 3.1, Mistral 7B, and NiuTrans LMT) and four NMT-based systems (Google Translate, Microsoft Translator, NLLB-200, and LibreTranslate-v1.5) in maintaining classifier-mediated sentiment consistency across twelve translation directions involving English, Spanish, French, and Chinese. Experiments were conducted on the Multilingual Amazon Reviews Corpus (MARC), comprising 84,000 randomly sampled user reviews. A multidimensional evaluation framework was used, combining sentiment-consistency metrics (accuracy, weighted F1, MCC, and SSR), translation-quality estimation (COMET-QE), and semantic-similarity assessment (LaBSE). Statistical significance was examined using the Friedman and Nemenyi post hoc tests. The results show that GPT-4o-mini, Gemini 2.5 Flash-Lite, and Google Translate consistently ranked among the strongest systems across multiple evaluation dimensions. Performance differences were particularly pronounced in translation directions involving Chinese, highlighting the influence of language-specific structural characteristics. Furthermore, semantic similarity and translation quality exhibited only moderate relationships with sentiment consistency, indicating that high semantic similarity does not necessarily guarantee strong sentiment consistency. Overall, the findings demonstrate the importance of multidimensional and statistically grounded evaluation frameworks for assessing cross-lingual sentiment consistency and provide practical insights into the strengths and limitations of contemporary MT systems.
E. Cetin, Çağrı Şahin· Applied Sciences· 0 citations
With the rapid development of artificial intelligence, AI translation tools have gained widespread popularity among Chinese university students in their daily lives and academic studies. College students have grown accustomed to relying on such tools to finish English writing and revise linguistic errors, making AI translation an indispensable auxiliary resource for their English learning. This paper selects three mainstream translation platforms, namely Doubao, Youdao Translate and ChatGPT, which represent domestic Chinese generative AI, traditional Chinese translation software and the most popular overseas generative AI respectively. This paper compares their translation accuracy on formal news articles and informal personal letters from three dimensions: grammar, lexical selection and textual coherence. Adopting data analysis and comparative analysis to ensure objectivity and rigor, this research intends to help undergraduates pick high-quality translation aids. The experimental results reveal that ChatGPT achieves the best overall translation performance, followed by Youdao Translate and then Doubao. In conclusion, ChatGPT is the most suitable option for Chinese college students in English writing practice.
Yitong Liao· Communications in Humanities...· 0 citations
Museum public signs mediate historical and cultural information for linguistically diverse audiences, and their translation requires both source-text accuracy and target-language accessibility. This qualitative case study applies Li Changshuan’s Comprehension-Expression-Adaptation (CEA) framework to nine Chinese-English bilingual exhibition-text pairs from the Nanning Museum of Urban Construction. The corpus comprises 43 photographs of bilingual signs collected during an on-site visit. Of the corresponding English translations, six exhibited comprehension-level problems, five expression-level problems, and three adaptation-level problems, while the remaining twenty-nine showed no noticeable problems. Comparative textual analysis was used to identify source-target discrepancies and to classify each case according to the CEA dimension that best represented the dominant problem, while recognising that individual examples may involve more than one dimension. The cases illustrate three recurring types of difficulty within the selected data: semantic misinterpretation and information loss at the comprehension level; grammatical, collocational, lexical, and tense-related problems at the expression level; and insufficient or inappropriate mediation of culture-specific and historical-institutional terms at the adaptation level. Revised translations are proposed for each case. On the basis of the case analysis and current museum-translation scholarship, the study recommends systematic verification of historical meaning and terminology, idiomatic target-language editing, and audience-oriented adaptation through selective explanation, transliteration, or reformulation where needed. Because the analysis is based on nine selected text pairs from a single museum, the findings are illustrative rather than estimates of the prevalence of particular error types. The study provides a focused application of the CEA framework to museum exhibition discourse and offers practical considerations for Chinese-English museum-text revision.
Unknown authors· Asian Journal of Education a...· 0 citations
This work introduces M-GATE (Multilingual Grammar, Accuracy in Translation, and Efficiency), a benchmark of linguistic proficiency spanning 30 typologically diverse languages from high- to low-resource, and evaluates over 50 models in more than 80 configurations.
Tomáš Burkert, Angelika Peljak-Łapińska, David Zelený· 0 citations
Functional evaluation proves more comprehensive than purely linguistic metrics in assessing translation quality in the AI-assisted translation era, and is operationalizing Nord's functionalist framework into a measurable evaluation model for systematically comparing HT and MT.
Dewi Rosnita Hardiany, M. F. R. Pratama, R. Nurjanah· Allure Journal· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.