Skip to content
Open access

SwitchEmbed: Representation Learning for Arabic-English Code-Switched Text

2026 · International Conference on Data Technologies and Applications · pp. 712-719 · 0 citations · 28 references
Computer Science

TL;DR

This work introduces CS-SNLI, a large-scale Arabic-English code-switched natural language inference dataset and adopts a contrastive learning framework with natural language inference supervision to encourage semantically consistent sentence embeddings across monolingual and code-switched variants.

Abstract

: Multilingual speakers often alternate between languages within a conversation, a phenomenon known as code-switching. This is common in Arabic-speaking communities, where Arabic and English are frequently mixed in everyday communication. Although recent advances in natural language processing have been driven by pretrained multilingual language models, these models are largely trained on monolingual data and often struggle to capture the abrupt language transitions and cross-lingual semantic interactions that characterize code-switched text. This work investigates representation learning for Arabic-English code-switched text at multiple levels. At the word level, we employ a code-switch-aware masked language modeling objective that captures token-level language variation and switch points. At the sentence level, we adopt a contrastive learning framework with natural language inference supervision to encourage semantically consistent sentence embeddings across monolingual and code-switched variants. To support this objective, we introduce CS-SNLI, a large-scale Arabic-English code-switched natural language inference dataset. The resulting embeddings are evaluated on sentiment analysis and named entity recognition. On sentiment analysis, the word-level pipeline improves F1-score by 2.17 percentage points over mBERT and 2.27 percentage points over XLM-R, while the sentence-level pipeline improves F1-score by 1.78 and 1.28 percentage points, respectively. In contrast, named entity recognition shows only marginal, statistically insignificant gains.

Read PDF

Similar papers

Preprint Aug 2026

AfriSwitch: A Benchmark for In-the-Wild African Code-Switched Speech Recognition

Code-switching is pervasive in bilingual African conversation, yet most ASR systems assume monolingual input and are evaluated on curated monolingual benchmarks. We present AfriSwitch, a 61.36-hour human-transcribed benchmark of in-the-wild code-switched speech spanning 16 African languages and language varieties, rele...

Gabrial Zencha Ashungafac, Busayo Awobade, Tobi Olatunji · 0 citations
Open access Aug 2026

Study of deep learning cues for cross linguistic part of speech tagging in English– Malayalam code-mixed data

The first structured analysis of micro macro divergence and rare class behavior in English Malayalam code mixed POS tagging is provided, providing the first structured analysis of micro macro divergence and rare class behavior in English Malayalam code mixed POS tagging.

Parvathy Padmakumar, Shreya S. Nair, A. Prajisha et al. · 0 citations
2026

Efficient Adaptation of English Language Models for Morphologically Rich and Underrepresented Languages: The Case of Arabic

A resource-efficient adaptation of the English-pretrained ModernBERT for Arabic, employing continued pretraining on large Arabic corpora followed by lightweight head-only fine-tuning with a frozen encoder, demonstrating that modern English encoder architectures can be efficiently transferred to Arabic through language-...

Ahmed Samy Eldamaty, M. Abdelrahman, Mohamed Mostafa Ibrahim Elbehery et al. · 0 citations

Code-Switch Pretraining for Improved Cross-Lingual Alignment in Low-Resource Languages

This work proposes an approach that integrates code-switching directly into masked language model pretraining, and introduces a multiview probabilistic translation strategy that samples candidate translations based on alignment likelihoods, applying substitutions only to unmasked tokens.

Ruan Visser, Trienko L. Grobler, Marcel Dunaiski · 0 citations
#natural language process... Preprint Sep 2026

IndicTriMix: Developing Language Identification Datasets and Models for Tri-Language Code-Mixing

This work proposes two approaches of code-mixed generation using parallel sentences of three languages and fine-tune contextual transformer-based models MuRIL and XLM-RoBERTa best suited for Indian languages for language identification in code-mixed settings.

Pruthwik Mishra, Rudra H. Trivedi, Avi Patel et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.