Skip to content
Preprint

Figurative and Cultural Knowledge in LLMs: Investigating Cross-Domain Transfer through Fine-Tuning

Aug 2026 · 0 citations · 39 references
Computer Science

TL;DR

The findings suggest that the relationship between culture and figurative language, though conceptually natural, is not straightforwardly captured through fine-tuning alone, and that fine-tuning reinforces experiential cultural knowledge while destabilizing historically grounded factual knowledge.

Abstract

Figurative language is deeply culturally embedded; fluent use requires not just linguistic competence but cultural immersion. We ask whether LLMs can learn this link: does fine-tuning on cultural data improve figurative language understanding, and vice versa? We conduct a systematic study across four models (ALLaM-7B, Fanar-1-9B, Qwen3-8B, Llama-3.1-8B) and six Arabic datasets spanning cultural commonsense, proverbs, and poetry across diverse dialects and regions. Fine-tuning on poetry improves idiom comprehension (+2.33%, p<0.05), a gain our ArabicMMLU control does not reproduce, indicating that it stems from figurative content rather than Arabic language adaptation and pointing to a sensitivity to non-literal meaning that transfers across figurative types. Cultural fine-tuning, by contrast, lowers proverb-interpretation accuracy in both Arabic-centric models. Transfer between the two domains is otherwise indistinguishable from noise, with Arabic models frequently regressing after fine-tuning, suggesting prior saturation of relevant knowledge, while multilingual models show greater adaptation headroom. Error analysis further reveals that fine-tuning reinforces experiential cultural knowledge while destabilizing historically grounded factual knowledge. Our findings suggest that the relationship between culture and figurative language, though conceptually natural, is not straightforwardly captured through fine-tuning alone.

View source

Similar papers

Review Open access Aug 2026

Artificial Minds, Cultural Shadows: Cultural Alignment, Identity, and Voice Across Multiple Large Language Models

Comparison of five widely used large language models suggests that AI-generated language may shape how culturally situated perspectives are expressed, with differences across models indicating that AI-generated language may shape how culturally situated perspectives are expressed.

Ashkan Goudarzi, Aylar Naderi Zonouz · 0 citations
#artificial intelligence Preprint Sep 2026

CulturalMenuBench: Probing the Knowledge-Application Gap in Multimodal Culinary Reasoning

Multimodal language models achieve near-ceiling scores on food recognition benchmarks, yet it remains unclear whether this success reflects genuine cultural understanding or mere visual matching. To probe this distinction, we introduce CulturalMenuBench, a benchmark of 4,870 items in 10 languages across 18 regions; its 10 tasks pair final-dish and step-by-step cooking images with ingredients, procedural text, and regional labels, spanning basic recognition to process-grounded cultural attribution. Evaluating 12 models exposes a substantial knowledge-application gap: models exceeding 94% on standard multiple-choice tasks drop to at most 56% when attributing dishes to Chinese regional cuisines, despite an identical four-way format. Diagnostic analyses explain why: error patterns are consistent with random guessing, accuracy tracks visual distinctiveness rather than cultural structure, and models classify cuisines more accurately from dish names alone than from images (+7-18 points). The knowledge is thus present but cannot be activated through visual input. An ablation confirms these tasks genuinely require procedural evidence: removing sequential cooking images selectively degrades process-grounded tasks while others remain stable. Overall, CulturalMenuBench shows that near-perfect recognition can conceal an inability to apply cultural knowledge, motivating training that explicitly connects perception, procedure, and cultural context. Code and data are publicly available.

Bo Zeng, Lin-Feng Gao, Pei-Qing Lin et al. · 0 citations
Open access Aug 2026

Cross-Cultural Scenario Benchmark: Evaluating LLMs’ Cross-Cultural Understanding

Cross-cultural reasoning and alignment have been identified as key weaknesses of large language models (LLMs), but the architectural or cognitive features underlying these failures have not been adequately examined. In addition, previous studies rely almost exclusively on datasets and benchmarks constructed under the WEIRD (Western, Educated, Industrialized, Rich, and Democratic) bias. To address this data bias issue, we prepare a dataset with substantial coverage of non-WEIRD cultures and five-dimensional (W, E, I, R, and D) annotations. This dataset supports an interpretable approach to examining weaknesses in LLMs’ cross-cultural alignment. We adopt the Chinese–Foreign Cultural Differences Case Repository at Xiamen University, which contains 9342 cases across 151 countries, 6 continents, and 10 cultural domains. These cases are processed and transformed into benchmark-ready structured data through topic normalization, structured metadata cleaning, continent correction, and country-level WEIRD annotation along five dimensions. Each case is converted into a six-option cultural attribution question with five cognitive-trap distractors grounded in cognitive reasoning and pragmatic interpretation. Evaluation of 6 mainstream large language models shows that their dominant failures do not involve explicit stereotypes. Instead, 61% of all errors arise from oversimplifying complex cultural phenomena or applying familiar cultural frames. The proportion of errors that explain specific cultural conflicts through seemingly universal value frames increases from 11% at the low-WEIRD end to 20% at the high-WEIRD end of the dataset. These results suggest that WEIRD data bias reflects both the underrepresentation of low-WEIRD cultures and the overactivation of dominant value frames in high-WEIRD contexts.

Meng-Xi Guo, Lei-Ming Gao, W. Zeng et al. · 0 citations
Preprint Aug 2026

Wisdom in Unity: The Role of Multilingual Training in Figurative Language Identification in Proverbs

Although multilingual approaches to figurative language identification are not new, the shift beyond language homogeneous training data requires a clearer understanding of the contribution of translated multilingual supervision. We examine this question using 742 proverb concepts across 6,787 translated instances in seven languages. We evaluate five models, including multilingual encoders and instruction tuned LLMs, under progressively increasing levels of multilingual supervision. Moreover, we introduce a multidimensional annotation framework for proverbs that characterizes them through four complementary figurative forms: Metaphorical, Moral/Advisory, Cause-Effect, and Culture Specific. Our findings show that approximately 50% of the translated multilingual training data is sufficient to achieve near-optimal figurative language identification performance. We further show that combining diverse figurative forms yields the strongest overall performance. A notable finding is that the least frequent figurative form, Culture Specific, exhibits the largest performance gains under multilingual supervision. Furthermore, the Moral/Advisory and Culture Specific forms contribute most to the performance of instruction-tuned LLMs on figurative language identification. These findings motivate multilingual figurative language identification to move beyond metaphor-centric taxonomies toward concept level multidimensional frameworks that explicitly model complementary forms of figurative meaning.

Rama Alomair, Remas Alsubaie, Walaa Saifalislam et al. · 0 citations
Preprint Aug 2026

Cultural Awareness is Represented but Not Decoded: Tracing Mythological Knowledge across 18 Open-Source LLMs

Open-source LLMs reliably name Zeus, Jupiter, and Thor, but recover their counterparts in less-represented traditions like Finnish, Slavic, Egyptian, or Chinese mythology far less consistently. We ask where inside the model this cultural default is produced. On a parallel cross-cultural substrate of Thompson-motif entities, we instrument 18 open-source LLMs from 8 architecture families with linear probing, logit lens, activation patching, and output extraction. The residual stream cleanly distinguishes cultures, well above a name-string baseline, yet the decoder collapses culturally-specific tokens onto dominant-tradition ones. The failure is at readout, not at representation. Asking the same question in the target culture's native language versus English produces failures that cluster within language but decouple across language: the decoder is gated on prompt language. We release a per-entity (probe, output) decomposition framework, a citation-anchored cross-cultural ground truth, a within- versus cross-mode correlation test for language-conditioned readout, and per-entity predictions for all 18 models.

Iaroslav Chelombitko, Ekaterina Chelombitko, Mika K. Hämäläinen · 0 citations
Open access Jul 2026

Reframing collocational errors as cultural collocational transfer: Insights from a Thai EFL learner corpus

This study investigates how culturally grounded conceptualizations shape English collocational usage in learner writing. Drawing on a 4.6-million-word corpus of Thai English as a foreign language (EFL) academic texts, collocations were extracted using corpus-driven association measures, including Mutual Information (MI) and t-score, and were manually validated against the British National Corpus (BNC). Of the 2,239 collocation types identified, 316 types (20,173 tokens) were classified through manual coding as cultural-transfer or culturally interpretable transfer-related forms when Thai linguistic patterning was accompanied by culturally salient conceptual motivation. Rather than viewing non-standard collocations as deficiencies, this study interprets this culturally classified subset as evidence of cultural meaning-making and lexical nativization. Quantitative and qualitative analyses revealed five domains: Kinship Hierarchy, Religious Practice, Food/Domestic Culture, Localized Service Economy, and Institutional/Media Discourse. Together, these domains accounted for 41% of all non-standard collocations, indicating that cultural transfer represents a systematic rather than incidental phenomenon in Thai learner English. The study contributes to applied corpus linguistics and World Englishes by showing how local cultural schemas are encoded in collocational choices. Pedagogically, the findings highlight the need to foster learners’ intercultural lexical awareness and to develop corpus-informed instructional practices that help learners negotiate cultural identity while maintaining communicative clarity.

Songtham Vongvirulh, A. Khamkhien · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.