Large language models (LLMs) increasingly evaluate human writing in high-stakes domains such as hiring and academic assessment, putting non-native speakers at particular risk. Drawing on the language attitudes framework, we compared human and LLM evaluations of parallel L1- and L2-written Japanese emails on three dimensions: fluency, status, and solidarity. Japanese raters rated L2 texts significantly lower on all three dimensions, with a fluency gap roughly twice the size of the status and solidarity gaps. Six LLM judges reproduced the direction of this bias, and five reproduced its ordering across dimensions. The models diverged from humans in two ways: all understated the solidarity gap, the most socially grounded dimension, and all differentiated among learner L1 backgrounds where humans did not. LLM judges thus reproduce native speakers'language attitudes in a structured yet attenuated form, and the language attitudes framework offers a ready-made yardstick for auditing them beyond English.
J-PragEval-v0 is introduced, a minimal-pair benchmark isolating four such phenomena from surface fluency, and Pragmatic Representation Steering is specified, a parameter-free inference-time method that edits residual-stream activations along the class-mean-difference directions probing identifies.
Large Language Models (LLMs) are increasingly used for translation, yet their value depends on preserving meaning rather than producing fluent output. This study evaluates seven LLMs on Japanese–Croatian translation, a low-resource, typologically distant language pair. Using rubric-based human evaluation of adequacy, fluency, terminology, and register, we compare model performance. Results show a stable ranking: qwen3 performs best, followed by phi4 and gemma3, while qwen2 performs worst. Performance differences reflect structural reconstruction, particularly argument recovery, aspectual mapping, lexical precision, and register. Qualitative analysis also reveals limited differentiation within the South Slavic continuum and pragmatic inconsistencies. Although productivity effects were not measured, improved translation adequacy may reduce post-editing and verification effort.
The Cross-Lingual Comprehension Gap (CLCG) is defined as the reduction in response quality when the same content and question are presented in a target language rather than in English.
Comparison of five widely used large language models suggests that AI-generated language may shape how culturally situated perspectives are expressed, with differences across models indicating that AI-generated language may shape how culturally situated perspectives are expressed.
Ashkan Goudarzi, Aylar Naderi Zonouz· Digital Studies in Language...· 0 citations
While large language models (LLMs) are increasingly used across social domains, current bias evaluations often rely on demographic proxies such as names, pronouns, and social categories. Linguistic variety itself receives less attention as an evaluative variable. This study therefore uses sociolinguistic variety to examine whether three LLMs respond differently to semantically equivalent prompts in General American English and Irish English. Thirty prompt pairs were administered twice to GPT-5, Claude, and Gemini, yielding 360 responses. Across the 180 paired comparisons, 92.8% differed in word count, although the direction varied by model: Claude produced longer Irish English responses on average, whereas GPT-5 and Gemini produced shorter responses. Claude explicitly referenced Irish English features in 55.0% of its Irish English responses, compared with 0.0% for GPT-5 and 1.7% for Gemini. No explicit correction, refusal, or researcher-observed tone shift occurred in either condition. These findings show variety-conditioned differences in response behavior, but they do not by themselves establish discriminatory biasThe study supports linguistic variety as an additional dimension for LLM differential-behavior evaluation and identifies model- and feature-specific patterns that warrant further evaluation with independent raters and additional varieties.
Awad H. Alshehri, N. Jaballah· Journal of Intelligent Decis...· 0 citations
Real-world communication often requires pragmatic reasoning: interpreting meanings implied through context and cultural convention rather than stated literally. Existing pragmatic evaluation remains largely limited to English and high-resource languages, leaving Indic languages unexplored despite their linguistic and cultural diversity. We introduce VakyArth, the first pragmatic benchmark for Indic languages, designed as a diagnostic evaluation covering Hindi, Punjabi, Tamil, and Malayalam. VakyArth evaluates models across five phenomena: deixis, speech acts, implicature, social pragmatics, and coherence; through multiple-choice questions, natural language inference, and translation, with all items authored by native speakers. Across multilingual large language models (LLMs) of varying families and sizes, we find consistent failures on pragmatic meanings rooted in Indic linguistic and cultural conventions. Our analysis shows systematic differences across languages and tasks: MCQ accuracy exceeds NLI accuracy in all model-language combinations, translation performance does not reliably track pragmatic understanding, and Indo-Aryan languages show a translation advantage over Dravidian languages. We further show that automatic translation metrics can miss fluent but pragmatically unfaithful outputs, especially for implicature and deixis.
Usneek Singh, Poorvaja Veera, B. Kumar et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.