Camellia is introduced, a benchmark for evaluating entity-centric cultural biases in nine Asian languages, spanning six Asian cultures, and it is found that LLMs can struggle with context understanding in some Asian languages, creating performance gaps between cultures in entity extraction.
Tarek Naous, Anagha Savit, Carlos Rafael Catalan et al.· arXiv.org· 2 citations· ⚡1
Role-playing large language models (LLMs) are expected to adopt a character's style while also respecting that character's knowledge boundaries. Prior evaluations detect character hallucination but rarely distinguish whether errors arise from failure to recognize a boundary or from failure to comply despite recognition. We introduce CHARM, a multicultural benchmark of 40 real and fictional characters drawn from five cultural-linguistic regions, and validated by native reviewers. It probes two boundary types, Temporal (historical vs. modern) and Cross-Universe (entities outside a character's narrative or historical universe), using abstention-enabled multiple-choice questions. We propose a two-stage evaluation that separates Boundary-Awareness (explicit recognition that a query is out of scope) from Boundary-Compliance (abstention when answering concrete questions). Evaluations across six LLMs show that hallucination is driven predominantly by compliance failures. Models frequently acknowledge that a query lies outside the character's knowledge yet still provide factual, out-of-character answers. By re-posing the same questions to the target character, we confirm that a large fraction of these cases are verified parametric overrides; the model stores the relevant fact but fails to suppress it. We also observe systematic cultural variation in these failures, consistent with imbalances in how characters from different regions are represented in model knowledge.
Su-Hyun Han, Nahyeon Park, Gaeun Seo et al.· 0 citations
It is highlighted that effective cultural alignment requires context-conditional modeling rather than uniform debiasing, and a new direction for mitigating entity-centric cultural bias in LLMs is established.
CultureConverse is introduced, a scalable, multilingual simulation and evaluation harness for culturally grounded assistant dialogue that covers 10 East and Southeast Asian regions, 58 subgroup identities, and 7 domains and performance gains from fine-tuning on 27,860 high-quality CultureConverse-DS samples improve in-domain assistance and transfer out-of-domain to cultural MCQ and safety classification benchmarks.
Bryan Chen Zhengyu Tan, Weihua Zheng, Thong T. Doan et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.