2026· International Conference on Data Technologies and Applications· pp. 712-719· 0 citations· 28 references
Computer Science
TL;DR
This work introduces CS-SNLI, a large-scale Arabic-English code-switched natural language inference dataset and adopts a contrastive learning framework with natural language inference supervision to encourage semantically consistent sentence embeddings across monolingual and code-switched variants.
Abstract
: Multilingual speakers often alternate between languages within a conversation, a phenomenon known as code-switching. This is common in Arabic-speaking communities, where Arabic and English are frequently mixed in everyday communication. Although recent advances in natural language processing have been driven by pretrained multilingual language models, these models are largely trained on monolingual data and often struggle to capture the abrupt language transitions and cross-lingual semantic interactions that characterize code-switched text. This work investigates representation learning for Arabic-English code-switched text at multiple levels. At the word level, we employ a code-switch-aware masked language modeling objective that captures token-level language variation and switch points. At the sentence level, we adopt a contrastive learning framework with natural language inference supervision to encourage semantically consistent sentence embeddings across monolingual and code-switched variants. To support this objective, we introduce CS-SNLI, a large-scale Arabic-English code-switched natural language inference dataset. The resulting embeddings are evaluated on sentiment analysis and named entity recognition. On sentiment analysis, the word-level pipeline improves F1-score by 2.17 percentage points over mBERT and 2.27 percentage points over XLM-R, while the sentence-level pipeline improves F1-score by 1.78 and 1.28 percentage points, respectively. In contrast, named entity recognition shows only marginal, statistically insignificant gains.
It is found that training on code-switched data aligns the representations of parallel text, particularly across different scripts, and that this alignment persists through training on monolingual documents.
Dries Rooryck, Alex Cai, Yonatan Belinkov et al.· 0 citations
Code-switching is pervasive in bilingual African conversation, yet most ASR systems assume monolingual input and are evaluated on curated monolingual benchmarks. We present AfriSwitch, a 61.36-hour human-transcribed benchmark of in-the-wild code-switched speech spanning 16 African languages and language varieties, rele...
The first structured analysis of micro macro divergence and rare class behavior in English Malayalam code mixed POS tagging is provided, providing the first structured analysis of micro macro divergence and rare class behavior in English Malayalam code mixed POS tagging.
Parvathy Padmakumar, Shreya S. Nair, A. Prajisha et al.· Frontiers in Big Data· 0 citations
A resource-efficient adaptation of the English-pretrained ModernBERT for Arabic, employing continued pretraining on large Arabic corpora followed by lightweight head-only fine-tuning with a frozen encoder, demonstrating that modern English encoder architectures can be efficiently transferred to Arabic through language-...
Ahmed Samy Eldamaty, M. Abdelrahman, Mohamed Mostafa Ibrahim Elbehery et al.· International Conference on...· 0 citations
This work proposes an approach that integrates code-switching directly into masked language model pretraining, and introduces a multiview probabilistic translation strategy that samples candidate translations based on alignment likelihoods, applying substitutions only to unmasked tokens.
Ruan Visser, Trienko L. Grobler, Marcel Dunaiski· 0 citations
This work proposes two approaches of code-mixed generation using parallel sentences of three languages and fine-tune contextual transformer-based models MuRIL and XLM-RoBERTa best suited for Indian languages for language identification in code-mixed settings.
Pruthwik Mishra, Rudra H. Trivedi, Avi Patel et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.