Skip to content
Preprint

Human-LLM Alignment in Language Attitudes Toward Non-Native Japanese

Aug 2026 · 0 citations · 2 references
Computer Science

Abstract

Large language models (LLMs) increasingly evaluate human writing in high-stakes domains such as hiring and academic assessment, putting non-native speakers at particular risk. Drawing on the language attitudes framework, we compared human and LLM evaluations of parallel L1- and L2-written Japanese emails on three dimensions: fluency, status, and solidarity. Japanese raters rated L2 texts significantly lower on all three dimensions, with a fluency gap roughly twice the size of the status and solidarity gaps. Six LLM judges reproduced the direction of this bias, and five reproduced its ordering across dimensions. The models diverged from humans in two ways: all understated the solidarity gap, the most socially grounded dimension, and all differentiated among learner L1 backgrounds where humans did not. LLM judges thus reproduce native speakers'language attitudes in a structured yet attenuated form, and the language attitudes framework offers a ready-made yardstick for auditing them beyond English.

View source

Similar papers

Preprint Aug 2026

Interpretable Cross-Lingual Alignment in Small Language Models: Probing Cultural and Pragmatic Reasoning in Japanese-English Bilingual LLMs

J-PragEval-v0 is introduced, a minimal-pair benchmark isolating four such phenomena from surface fluency, and Pragmatic Representation Steering is specified, a parameter-free inference-time method that edits residual-stream activations along the class-mean-difference directions probing identifies.

F. Braun · 1 citation
Open access Jul 2026

Large Language Models for Japanese–Croatian Translation: Human Evaluation and Macroeconomic Implications

Large Language Models (LLMs) are increasingly used for translation, yet their value depends on preserving meaning rather than producing fluent output. This study evaluates seven LLMs on Japanese–Croatian translation, a low-resource, typologically distant language pair. Using rubric-based human evaluation of adequacy, fluency, terminology, and register, we compare model performance. Results show a stable ranking: qwen3 performs best, followed by phi4 and gemma3, while qwen2 performs worst. Performance differences reflect structural reconstruction, particularly argument recovery, aspectual mapping, lexical precision, and register. Qualitative analysis also reveals limited differentiation within the South Slavic continuum and pragmatic inconsistencies. Although productivity effects were not measured, improved translation adequacy may reduce post-editing and verification effort.

Ratomir Karlović, Mieta Bobanović Dasko, Irena Srdanović · 0 citations
Review Open access Aug 2026

Artificial Minds, Cultural Shadows: Cultural Alignment, Identity, and Voice Across Multiple Large Language Models

Comparison of five widely used large language models suggests that AI-generated language may shape how culturally situated perspectives are expressed, with differences across models indicating that AI-generated language may shape how culturally situated perspectives are expressed.

Ashkan Goudarzi, Aylar Naderi Zonouz · 0 citations
Open access Aug 2026

Linguistic Variety as a Framework for Evaluating LLM Behavior: A Comparative Study of Irish English and General American in Large Language Model Responses

While large language models (LLMs) are increasingly used across social domains, current bias evaluations often rely on demographic proxies such as names, pronouns, and social categories. Linguistic variety itself receives less attention as an evaluative variable. This study therefore uses sociolinguistic variety to examine whether three LLMs respond differently to semantically equivalent prompts in General American English and Irish English. Thirty prompt pairs were administered twice to GPT-5, Claude, and Gemini, yielding 360 responses. Across the 180 paired comparisons, 92.8% differed in word count, although the direction varied by model: Claude produced longer Irish English responses on average, whereas GPT-5 and Gemini produced shorter responses. Claude explicitly referenced Irish English features in 55.0% of its Irish English responses, compared with 0.0% for GPT-5 and 1.7% for Gemini. No explicit correction, refusal, or researcher-observed tone shift occurred in either condition. These findings show variety-conditioned differences in response behavior, but they do not by themselves establish discriminatory biasThe study supports linguistic variety as an additional dimension for LLM differential-behavior evaluation and identifies model- and feature-specific patterns that warrant further evaluation with independent raters and additional varieties.

Awad H. Alshehri, N. Jaballah · 0 citations
#artificial intelligence Preprint Sep 2026

VakyArth: Evaluating Pragmatic Competence in LLMs across Indic Languages

Real-world communication often requires pragmatic reasoning: interpreting meanings implied through context and cultural convention rather than stated literally. Existing pragmatic evaluation remains largely limited to English and high-resource languages, leaving Indic languages unexplored despite their linguistic and cultural diversity. We introduce VakyArth, the first pragmatic benchmark for Indic languages, designed as a diagnostic evaluation covering Hindi, Punjabi, Tamil, and Malayalam. VakyArth evaluates models across five phenomena: deixis, speech acts, implicature, social pragmatics, and coherence; through multiple-choice questions, natural language inference, and translation, with all items authored by native speakers. Across multilingual large language models (LLMs) of varying families and sizes, we find consistent failures on pragmatic meanings rooted in Indic linguistic and cultural conventions. Our analysis shows systematic differences across languages and tasks: MCQ accuracy exceeds NLI accuracy in all model-language combinations, translation performance does not reliably track pragmatic understanding, and Indo-Aryan languages show a translation advantage over Dravidian languages. We further show that automatic translation metrics can miss fluent but pragmatically unfaithful outputs, especially for implicature and deixis.

Usneek Singh, Poorvaja Veera, B. Kumar et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.