Aug 2026· International Journal of Bilingualism· 0 citations· 25 references
Abstract
Recently, contact linguistics has become increasingly interested in multiword units. At the same time, the code-copying framework (CCF) includes the notion of mixed copies (MCs) that are in-between global copies (‘borrowing’) and selective copies (‘structural change’) and illustrate the transition between the lexicon and grammar. The research question is: What types of MCs occur in English-Estonian bilingual speech?
The data were transcribed, and MCs identified, annotated, and classified according to their structure. English items were searched for in Estonian dictionaries to establish their Estonian equivalents or conventionalization of such items. The frequencies of MCs and their Estonian equivalents were also searched on Google to determine whether the MCs occur outside the corpus.
Three datasets were analysed: written texts from 44 blogs (385,124 tokens), spoken data from 10 vlogs (117,555 tokens), and 8 podcasts (77,277 tokens). Quantitative analyses of the various MC types were conducted, followed by a qualitative analysis of representative examples.
Compound nouns constitute the majority of MCs, followed by idioms, phrasal compounds, and a small number of compound verbs. No frame-changing MCs (i.e., MCs resulting in grammatical change) were attested. Since compound nouns and analytic verbs occur in both languages, structural similarity may be a facilitating factor in copying.
The notion of MCs is not widely used. Research typically focuses on particular types of items (e.g., compound nouns or verbs); here, however, the question is reversed: which types of items yield MCs?
It was established that the proportion of MCs in the data is comparable to that of selective copies. Within MCs, the globally copied element renders the remaining part more specific, highlighting the importance of meaning in contact-induced language change. MCs are also present on the Estonian internet and, in some cases, outnumber their Estonian equivalents, if such equivalents exist.
This study determines the status of English-origin nouns in otherwise Vietnamese discourse among bilingual speakers in Australia. It examines whether these nouns function as fully integrated loanwords, nonce borrowings, or instances of code-switching.
Taking a comparative variationist approach, this study compares patterns of syntactic integration across three points of comparison: Vietnamese nouns (VN), attested English-origin loanwords (LW) and English-origin single nouns (EN1) and multi-word constructions (EN2).
The study is based on a quantitative analysis of natural, spontaneous speech produced by nine bilingual speakers (five female, four male). The analysis compares patterns of post-nominal modification and quantifier and classifier usage across the points of comparison to disambiguate the status of English-origin nouns in Vietnamese discourse.
Findings reveal overall low levels of syntactic integration for EN1, suggesting that most of these items reflect nonce borrowing or code-switching rather than fully conventionalised forms. EN1 shows stronger integration into Vietnamese grammar, particularly in quantified contexts where classifier usage aligns with VN norms and in post-nominal possessives. EN2, however, exhibits minimal adaptation, retaining more English-like structures and showing no transparent patterns of code-switching, suggesting intermediary mixing strategies that blur the boundaries between borrowing and code-switching.
These findings contribute to language mixing theory by providing empirical evidence on the syntactic behaviour of English-origin nouns in otherwise Vietnamese discourse. The study also holds practical implications for heritage language maintenance by providing insights into how bilinguals manage grammar and vocabulary dynamically in multilingual settings.
This research provides a nuanced variationist analysis of English-origin nouns in Vietnamese discourse. It contributes to the study of mixing strategies between English and Vietnamese, an understudied language pair. Methodologically, this study adds to the growing body of studies that use quantitative and comparative methodology for identifying linguistic patterns.
Hoa Do, James A. Walker· International Journal of Bil...· 0 citations
Language models are often evaluated as though capabilities demonstrated in English remain equally available when the same content is presented in other languages. Traditional multilingual benchmarks rarely isolate language while holding content, question, reference answer, model, and evaluation unit constant. We define the Cross-Lingual Comprehension Gap (CLCG) as the reduction in response quality when the same content and question are presented in a target language rather than in English. Using ParallelQA-18, a professionally human-translated parallel corpus, we evaluate five models from five laboratories on a stratified sample of 150 articles across 18 languages (English reference; Portuguese high-resource baseline; 16 targets spanning Joshi et al. 2020 classes 0-4). A within-item design varies only passage language. The primary estimator contrasts English versus pooled target-language Token-F1 micro-means on higher-complexity open-ended questions, with article-cluster bootstrap intervals. The primary pooled CLCG is 0.078 (95% CI 0.072-0.084), about a 17% reduction relative to the English score; the equal-language macro summary is 0.077. Net of Portuguese, the macro gap is 0.016 (95% CI 0.013-0.020). Language-level CLCG is negatively associated with Joshi resource class (rho = -0.594, p = 0.015, n = 16). In blinded paired human evaluations, higher-resource responses are preferred in 61.6% of decisive judgments (estimated preference probability 0.655, 95% CI 0.558-0.741). Capabilities shown in English should not be assumed to transfer equally to other languages; English-centered evaluations may overestimate quality for users of low-resource languages.
Social media has become a major venue for multilingual communication, where users frequently mix multiple languages within a single utterance. Although code-mixed corpora have been developed for several language pairs, Persian-English code-mixing remains relatively underexplored. Existing Persian resources lack Universal Dependencies (UD) part-of-speech (POS) annotations for code-mixed words, limiting both linguistic analyses and the development of syntax-aware NLP models. To address this gap, we introduce PERCEPT, the first publicly available large-scale Persian-English code-mixed corpus annotated with Universal Dependencies POS tags for code-mixed words. The dataset comprises 6,800 posts collected from X, Instagram, and Digikala. We further present an LLM-assisted annotation framework that automatically assigns POS tags and document-level topics. Human evaluation demonstrates high agreement between the automatically generated annotations and gold annotations, confirming the reliability of the annotations. Using PERCEPT, we conduct the first comprehensive linguistic analysis of Persian-English code-mixing across multiple social media platforms. Our analyses reveal that nouns are the predominant category for code-mixed words, while the distributions of other POS categories vary across platforms. We further find that the positional distribution of code-mixed words is remarkably consistent across platforms, whereas the triggering effect is substantially more pronounced in Digikala. PERCEPT is publicly available at https://github.com/kalhorghazal/PERCEPT.
Ghazal Kalhor, Zahra Jafari, Amirarsalan Shahbazi et al.· 0 citations
This study investigates morphological code-mixing, the combination of morphemes from two languages within a single word, in the spontaneous speech of young German–English bilingual children. Following a usage-based perspective, the study asks whether children’s morphological code-mixing shows systematic patterns and whether these patterns reflect properties of the linguistic input they receive.
The analysis is based on longitudinal corpus data from three German–English simultaneous bilingual children aged 2 to 3 years. All children were raised in Germany in families following a one-parent–one-language approach, but their input conditions differed in the relative frequency and distribution of German and English. Utterances were coded for language type and analyzed for instances of morphological code-mixing.
Across the three corpora, 696 instances of morphological code-mixing were identified. Each mixed form was coded for type of mixing, part of speech, language of stems and affixes, and grammatical categories realized in inflectional mixing. Quantitative comparisons were conducted across children to examine distributional patterns and effects of differing input conditions.
Morphological code-mixing appeared in fewer than 2% of utterances but was produced by all children. Individual differences emerged in both frequency and types of mixes. Inflectional mixing was the most common pattern, especially German inflectional morphemes attached to English stems. Children with more balanced or shifting input conditions also produced German stems with English inflectional markers. Mixed compounds occurred in both German–English and English–German configurations. These patterns mirrored the children’s input distributions and developmental shifts in language proficiency.
The results support a usage-based account in which morphological code-mixing arises from entrenched morphological frames and analogical extension rather than violations of grammatical constraints.
The study shows that morphological code-mixing in early bilingual development is systematic, input-sensitive, and shaped by children’s experience with partially schematic constructions across their two languages.
Antje Endesfelder Quick, Stefan Hartmann, Nikolas Koch· International Journal of Bil...· 0 citations
Significant advances have been achieved in machine translation (MT) in recent times, particularly state of the art (SOTA) models for languages like English and Indian having distinct grammatical structures and limited monolingual training data. This paper analyses various recent state-of-the-art variants of large language models (LLMs) and neural machine translation (NMT) for Indian languages in comparison to statistical machine translation (SMT). It tackles key questions, such as idiomatic expressions, morphologically complex grammar or the scarceness of parallel corpora. Furthermore, it studies bytewise BPE, compares translation models in terms of BLEU scores using separate and shared-vocabulary representation with copy actions between the BPE translations, and analyses how multitask learning (Caruana (1997)) and attention mechanisms can contribute to the quality of translation. In summary, it provides directions for future work by suggesting new avenues of research including better curated datasets, more efficient approaches for lowresource languages and culturally aware translations.
Jayanand A. Kamble, S. Jadhav, V. J. Kadam· International Journal of Inf...· 0 citations
Old Tamil literature differs completely from Modern Tamil in both word structure and sentence structure. Tense markers or infixes play a crucial role in denoting the tense of verbs. Although Tolkappiyar defined the tenses as three, he did not analyze and specify the exact tense markers for them. It was Nannular, who came later, who categorized the four markers—-t- (-த்-), -ṭ- (-ட்-), -ṟ- (-ற்-), and -in- (-இன்-)—as past tense infixes.
Beyond these boundaries, various allomorphs and morphophonemic variations were in practice in Old Tamil. This paper comprehensively examines past tense markers using evidence from Sangam literature, grounded in the theories of Descriptive Linguistics, Historical Comparative Grammar, and Morphophonemics.
In particular, it categorizes phonologically conditioned allomorphs, final consonant doubling (gemination), and other past tense markers not mentioned by traditional grammarians (-tt-, -nt-, -i-, -n-, -y-, -k-) from a morphological perspective, providing a detailed explanation of their historical background and universal linguistic parallels.
முனைவர் நா. சரண்யா· Tamilmanam International Res...· 0 citations