TanglishGuard: Benchmarking AI Safety Guardrails on Tamil-English Code-Mixed Prompts
Abstract
Large Language Models (LLMs) have become indispensable across modern artificial intelligence applications, yet their safety mechanisms continue to be evaluated almost exclusively on English inputs. In multilingual nations like India, where millions communicate daily through Tanglish a fluid blend of Tamil and English this evaluation gap raises fundamental concerns about whether these systems respond safely and consistently across diverse linguistic contexts. This study addresses this critical oversight by introducing TanglishGuard, a novel benchmark designed to systematically assess the robustness of LLM safety mechanisms against code-mixed inputs. The benchmark evaluates four state-of-the-art models ChatGPT, Gemini, Claude, and DeepSeek using equivalent harmful prompts expressed in English, Tamil, and Tanglish across nine distinct harm categories. Through rigorous experimentation, our findings reveal that while all models demonstrate strong safety compliance on English and Tamil inputs, Tanglish prompts reveal subtle but consistent vulnerabilities. These inconsistencies manifest as occasional failures in detecting harmful intent within code-mixed language, highlighting significant gaps in multilingual AI safety frameworks. TanglishGuard provides a practical, reproducible framework for evaluating safety in mixed-language settings, offering empirical evidence that current safety evaluations are insufficient for real-world multilingual usage. The benchmark contributes to the development of more robust and equitable AI systems by ensuring safety mechanisms are tested against authentic communication patterns rather than sanitised English-only datasets. This work underscores the urgent imperative to move beyond English-centric safety evaluations. As AI systems become increasingly embedded in diverse linguistic communities worldwide, ensuring their safety across the full spectrum of human language use is not merely a technical challenge but a fundamental requirement for fairness, equity, and responsible AI deployment.