A Comprehensive Survey on Code-Switched LLM Safety and Robustness Evaluation
Abstract
Large language models (LLMs) have demonstrated remarkable capabilities across languages, yet their safety alignment remains predominantly evaluated in monolingual, especially English, settings. Code-switching (also known as codemixing)—the alternation between two or more languages within a single utterance or conversation—is a pervasive phenomenon in multilingual societies and digital communication. Recent evidence reveals that code-switched inputs can substantially degrade the safety robustness of LLMs, enabling jailbreaks and harmful outputs that monolingual safety mechanisms fail to intercept. This survey provides the first comprehensive review of research on code-switched LLM safety and robustness evaluation. We systematize existing attack methodologies (including Code-Switching Red-Teaming, Multilingual Blending, and attributional analyses), evaluation datasets and metrics, empirical findings on attack success rates and linguistic factors, and emerging mitigation strategies. We further situate these works within the broader multilingual safety literature, highlight critical gaps in linguistic coverage, cultural contextualization, and mechanistic understanding, and outline a research agenda toward equitable, linguistically inclusive LLM safety. Our synthesis aims to guide researchers and practitioners in developing more robust evaluation frameworks and alignment techniques for real-world multilingual deployments.