Skip to content
Open access

The Corpus of English–Tagalog Code-switching: an integrative corpus of linguistic contact effects

Sep 2026 · Scientific Data · 0 citations

Abstract

The Corpus of English–Tagalog Code-switching (CEnTaCS) is a new corpus designed to document naturalistic English–Tagalog bilingual speech. Locally referred to as Taglish, it is a widely-used yet understudied contact variety in the Philippines. Recorded in 2025 in Metro Manila, the corpus comprises three interrelated datasets: individual story retelling recordings, cognitive control data based on the Arrow Flanker task, and sociolinguistic questionnaire data. CEnTaCS offers fine-grained data on code-switching across clausal, lexical, and morphemic levels, with particular attention to intra-word code-switching, a salient but still insufficiently documented feature of Taglish. Beyond code-switching, the corpus captures a broad range of language contact phenomena, including borrowing, calquing, convergence, and interference. By making these data systematically available, the corpus supports an integrated account of mixed-language practices that brings together structural, cognitive, and sociolinguistic perspectives. CEnTaCS thus enables researchers to investigate how contact-induced structures are constrained linguistically, how they relate to bilingual processing, and how they index social meaning and identity in a community where bilingualism is deeply embedded in everyday life.

Read PDF

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.