SACMQ- South Asian Cultural Multilingual Question Dataset
Abstract
SACMQ — South Asian Cultural Multilingual Question Dataset contains multilingual question-answering data in five South Asian languages: Bangla, Hindi, Urdu, Tamil, and Nepali. It was developed to evaluate large language models (LLMs) and Retrieval-Augmented Generation (RAG) systems on multilingual knowledge, with topical cluster labels and an is_cultural field supporting comparisons between culturally grounded and non-cultural questions. The dataset was constructed from titles and contextual passages collected from the corresponding regional Wikipedia editions. For each retained record, a two-line contextual excerpt was derived and supplied to Qwen 3 32B to generate a question whose intended answer is the corresponding article title. The generated questions are preserved as experimental model outputs and may therefore contain ambiguity, title leakage, duplicate wording, punctuation variation, or other generation artifacts. Each record includes the native-language title, translated title fields where applicable, full context, two-line context, generated question, topical cluster label, and is_cultural label. The topical clusters include economics, entertainment, event, food, health, historical, language, locations, material, miscellaneous, organization, people, politics, religion, sports, and temporal. The is_cultural field identifies whether a record is classified as culturally grounded or non-cultural, enabling evaluation across cultural categories alongside language and topic. The dataset contains 35,736 records in total: 10,278 Bangla, 8,694 Hindi, 5,806 Urdu, 8,553 Tamil, and 2,405 Nepali entries. SACMQ was developed as part of the research project “Analyzing Performance of Retrieval Augmented Generation Models for QA Tasks” at the Department of Computer Science and Engineering, Brac University.