Skip to content

Feeding BabyLMs Macaroni: Code-Switching Curricula Cause Cross-Lingual Convergence

Sep 2026 · 0 citations · 68 references
Computer Science

TL;DR

It is found that training on code-switched data aligns the representations of parallel text, particularly across different scripts, and that this alignment persists through training on monolingual documents.

Abstract

Children in multilingual communities often code-switch, using multiple languages in a single utterance. Can we induce cross-lingual alignment in language models by training on code-switched text? We pretrain small decoder-only transformers on two 100M-word multilingual corpora: a base corpus formed by mixing the English, Dutch, and Chinese BabyBabelLM datasets, and a corpus generated from it by inserting word- and sentence-level code-switching using an LLM. We find that training on code-switched data aligns the representations of parallel text, particularly across different scripts, and that this alignment persists through training on monolingual documents. Under a learning curriculum that progresses from word-level code-switching, to sentence-level code-switching, to monolingual documents, models trained on code-switched data outperform baselines trained without it on the BabyLM evaluation suite. Our work characterizes code-switching curriculum learning as an effective data augmentation method for multilingual pretraining. We release our code, data, and models at https://github.com/drooryck/multilingual-macaroni.

View source

Similar papers

Preprint Aug 2026

AfriSwitch: A Benchmark for In-the-Wild African Code-Switched Speech Recognition

Code-switching is pervasive in bilingual African conversation, yet most ASR systems assume monolingual input and are evaluated on curated monolingual benchmarks. We present AfriSwitch, a 61.36-hour human-transcribed benchmark of in-the-wild code-switched speech spanning 16 African languages and language varieties, rele...

Gabrial Zencha Ashungafac, Busayo Awobade, Tobi Olatunji · 0 citations

Code-Switch Pretraining for Improved Cross-Lingual Alignment in Low-Resource Languages

This work proposes an approach that integrates code-switching directly into masked language model pretraining, and introduces a multiview probabilistic translation strategy that samples candidate translations based on alignment likelihoods, applying substitutions only to unmasked tokens.

Ruan Visser, Trienko L. Grobler, Marcel Dunaiski · 0 citations
#natural language process... Preprint Sep 2026

Language Discrimination Improves Linguistic Learning in Multilingual Speech Models

Multilingual self-supervised speech models can benefit from sharing information across languages, but under a matched total pretraining data budget they still fall short of monolingual models. We show that strengthening the model's ability to discriminate languages during pretraining reduces and, on some measures, clos...

Maureen de Seyssel, Jie Chi, Zakaria Aldeneh · 0 citations
#artificial intelligence Preprint Sep 2026

Why Pretraining Fails to Share Cross-Lingual Knowledge

Large Language Models (LLMs) have made remarkable progress in the processing and modeling of many languages. Yet, unlike human multilinguals, they exhibit surprisingly limited cross-lingual knowledge transfer. While this limitation is well documented, its origins during multilingual training remain unclear. We pretrain...

Adam Gaber, Uriel Dolev, Elisabeth Fittschen et al. · 0 citations
#natural language process... Preprint Sep 2026

Beetle: A Bilingual Model Suite for Modelling Second-Language Processing

Bilingual language models (LMs) offer a controlled setting for studying how training conditions shape second-language (L2) behaviour, but prior work typically varies exposure structure, scale, and architecture at once, making it difficult to attribute effects to any single factor. We introduce Beetle, a controlled lang...

Suchir Salhan, Catherine Arnett, James A. Michaelov et al. · 0 citations
#natural language process... Preprint Sep 2026

IndicTriMix: Developing Language Identification Datasets and Models for Tri-Language Code-Mixing

This work proposes two approaches of code-mixed generation using parallel sentences of three languages and fine-tune contextual transformer-based models MuRIL and XLM-RoBERTa best suited for Indian languages for language identification in code-mixed settings.

Pruthwik Mishra, Rudra H. Trivedi, Avi Patel et al. · 0 citations

Related blog posts

MIT News · Artificial Intelligence Sep 24, 2026

Estimating suicide risk from text

A new language-processing tool could help identify the highest-risk individuals from natural language, enabling swifter interventions.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.