Aug 2026· Tamilmanam International Research Journal of Tamil Studies· Vol 12, pp. 609-627· 0 citations
TL;DR
A small-scale pilot Kongu Tamil benchmark dataset was curated by the author based on a verified Kongu dialect vocabulary and preliminary findings indicate that while the language model handles lexical dialectal variations reasonably well, significant errors occur in cultural nuances, syntactic variations, and idiomatic expressions.
Abstract
While Large Language Models (LLMs) demonstrate significant capabilities in standard Tamil, the extent to which they accurately understand and translate Kongu regional Tamil—spoken in the Kongu region including Coimbatore, Tiruppur, Erode, Salem, and Karur—remains an insufficiently explored area. This article examines the dialect-handling capabilities of AI language models, focusing on Kongu Tamil. For this purpose, a small-scale pilot Kongu Tamil benchmark dataset was curated by the author based on a verified Kongu dialect vocabulary. Using this dataset, zero-shot and few-shot experiments were conducted with a publicly accessible Large Language Model (Claude), and the responses were evaluated based on the author's linguistic knowledge, with errors categorized accordingly. Additionally, citing results from previously published studies comparing various LLMs (Gemini, ChatGPT, Claude), this paper proposes a comprehensive methodological plan for a complete multi-model comparative study for Kongu Tamil, along with a human expert evaluation protocol. Preliminary findings indicate that while the language model handles lexical dialectal variations reasonably well, significant errors occur in cultural nuances, syntactic variations, and idiomatic expressions. This study underscores the need to build large-scale, human-expert-verified benchmark datasets for Tamil dialects.
With the rapid growth of computer technology, the application of the Tamil language has expanded globally. Overcoming the early limitations of 8-bit ASCII encoding systems, the introduction of 16-bit encoding standards such as Unicode and TACE16 has enabled the seamless use of Tamil characters across all computer and m...
In the second decade of the twenty-first century, Artificial Intelligence (AI) technology is creating a profound impact across various domains, including language, culture, media, and commerce. This technology has unlocked new possibilities to globally monetize, disseminate, and preserve Tamil language and cultural con...
V. K. Veerakumar, G. Sasikaladevi, D. Rajakumari· Tamilmanam International Res...· 0 citations
MameLoshnLM, the first open-source 8B-parameter language model built specifically for Yiddish, is presented, providing both a foundation for Yiddish NLP and a practical template for language model development in historically rich but digitally underrepresented languages.
Uri Katz, Omer Goldman, Tomasz Limisiewicz et al.· 0 citations
Cantonese is widely spoken but remains low-resource in written data, with no large corpus of native Cantonese reasoning traces available for model training. We develop and release CantoneseLLM v2, comprising models based on Qwen3 8B and 30B-A3B. The models are trained through CPT on 784 million Cantonese and Hong Kong-...
Tsz-Chung Cheng, Chung-Shing Cheng, Chaak-ming Lau et al.· 0 citations
Perumal Murugan stands out as a writer who has deeply documented the regional life and linguistic style of the Kongu region in modern Tamil literature. Across his nearly thirty years of literary output—particularly in the novels Koolamadhari, Madhorubagan, Aalandapatchi, and Pookuzhi—Kongu dialect usages, agriculture-r...
முனைவர் இல. பூவலிங்கம்· Tamilmanam International Res...· 0 citations