Skip to content

Code-Switch Pretraining for Improved Cross-Lingual Alignment in Low-Resource Languages

· 0 citations · 26 references

TL;DR

This work proposes an approach that integrates code-switching directly into masked language model pretraining, and introduces a multiview probabilistic translation strategy that samples candidate translations based on alignment likelihoods, applying substitutions only to unmasked tokens.

View source

Similar papers

#artificial intelligence Preprint Sep 2026

Why Pretraining Fails to Share Cross-Lingual Knowledge

Large Language Models (LLMs) have made remarkable progress in the processing and modeling of many languages. Yet, unlike human multilinguals, they exhibit surprisingly limited cross-lingual knowledge transfer. While this limitation is well documented, its origins during multilingual training remain unclear. We pretrain...

Adam Gaber, Uriel Dolev, Elisabeth Fittschen et al. · 0 citations
#natural language process... Preprint Sep 2026

Cross-Lingual Representation Alignment by Token-Level Optimal Transport in a Language-Agnostic Space

Cross-lingual alignment (CLA) aims to align the representations of large language models (LLMs) across languages, enabling cross-lingual transfer to improve multilingual capabilities. Previous CLA methods often ignore language-specific information encoded in representations and only consider sentence-level alignment, w...

Taisei Yamamoto, Ryoma Kumon, D. Bollegala et al. · 0 citations
#natural language process... Preprint Sep 2026

Sequential Adapter Stacking for Cross-Lingual Low-Resource ASR

Extending large-scale multilingual automatic speech recognition (ASR) models to low-resource languages remains challenging. Model performance is skewed toward high-resource languages and degrades sharply for languages with limited labeled data and pre-training exposure. To address this, we investigate parameter-efficie...

Thai Thi Thanh Thao Dang, Meng-Jie Qian, Kate Knill · 1 citation
#natural language process... Preprint Aug 2026

One Form to Transfer Them All: Pretraining Multilingual Language Models Beyond Native Orthography

These results indicate that for multilingual models spanning typologically diverse scripts, to obtain maximum benefits, romanization should be treated as a core design choice applied at pretraining rather than a post hoc fix.

Mu-Ge Zhang, Aaron Jencks, Krishna Badikela et al. · 0 citations
Preprint Aug 2026

Predicting Multilingual Classification and Translation Performance of LLMs with Cross-Lingual Alignment -- Is English Enough?

A PMI-based translation metric is proposed, which is less dependent on the target language and correlates strongly with chrF, and finds that CLA with English predicts translation quality comparably to or better than source-target CLA.

Adnan Al Ali, Kathy Hämmerl, Jindrich Libovický et al. · 0 citations
Preprint Aug 2026

Scaling Unsupervised Word Alignment to Documents via Structural Constraints

CTFAlign is introduced, a lightweight, training-free approach for document-level word alignment that applies a coarse-to-fine refinement strategy that restricts the alignment search space to semantically similar regions and introduces MDPAlign, a simpler alternative that constrains alignments by position with a main di...

Michelle Wastl, Jannis Vamvas, Rico Sennrich · 1 citation

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.