Current cultural evaluations for large language models (LLMs) often reduce culture to single-turn factual recall via MCQs, failing to capture a common use case: users seeking practical help over multiple turns in culturally grounded scenarios. We introduce CultureConverse, a scalable, multilingual simulation and evalua...
Bryan Chen Zhengyu Tan, Wei-Hua Zheng, Thong T. Doan et al.· 0 citations
This work describes the near-optimal region, the set of allocations within a specified tolerance of peak performance, which is wide even for small tolerances, widens with model scale, and transfers reliably from small proxy models to large target models.
Jingtan Wang, A. Verma, Xiaoqiang Lin et al.· 0 citations
CultureConverse is introduced, a scalable, multilingual simulation and evaluation harness for culturally grounded assistant dialogue that covers 10 East and Southeast Asian regions, 58 subgroup identities, and 7 domains and performance gains from fine-tuning on 27,860 high-quality CultureConverse-DS samples improve in-...
Bryan Chen Zhengyu Tan, Weihua Zheng, Thong T. Doan et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.