Skip to content
Open access

Study of deep learning cues for cross linguistic part of speech tagging in English– Malayalam code-mixed data

Aug 2026 · Frontiers in Big Data · Vol 9 · 0 citations · 18 references
Medicine

TL;DR

The first structured analysis of micro macro divergence and rare class behavior in English Malayalam code mixed POS tagging is provided, providing the first structured analysis of micro macro divergence and rare class behavior in English Malayalam code mixed POS tagging.

Abstract

Introduction Part Of Speech (POS) tagging is a fundamental task in Natural Language Processing (NLP) that assigns grammatical labels to words in a sentence. Code mixed text, which entails switching between two or more languages within a single conversation or a sentence, presents challenges for POS tagging. This investigation entailed a comprehensive study of deep learning approaches for cross linguistic POS tagging, focused on English Malayalam code mixed data prevalent on social media platforms. The study was carried out on linguistically complex and varied English Malayalam code mixed text from social media platforms with informal spellings, language switching, transliteration, slang, and unclear grammatical boundaries, reflecting the characteristics of informal online communication. We provide the first structured analysis of micro macro divergence and rare class behavior in English Malayalam code mixed POS tagging. Methods We evaluated 14 state of the art model configurations that span traditional sequence labeling approaches and multilingual transformer architectures. Models were compared using standard performance metrics prevalent in the domain of data science, supplemented by normalized confusion matrices, error prone tag identification and micro-macro F1 gap analysis. Results Our results showed that CRF (No Lang) emerged as the most balanced model overall on macro F1 (all classes) of 0.8170, while (BiLSTM + CRF) achieved the highest macro F1 (seen classes) of 0.8831, precision of 0.9167, and recall of 0.875, though this reflects strong performance concentrated on frequent tag classes rather than balanced coverage across the full tag set. Notably, the pretrained multilingual transformers (mBERT, MuRIL), despite prior exposure to Malayalam during pretraining, were outperformed on several key metrics by CRF and BiLSTM models trained directly on the code mixed dataset. Discussion This finding was contrary to our expectation that existing multilingual knowledge would translate into a clear advantage on this task and merits further investigation.

Read PDF

Similar papers

Preprint Aug 2026

PERCEPT: A Corpus for POS Tagging and Analysis of Persian-English Code-Mixing

This work introduces PERCEPT, the first publicly available large-scale Persian-English code-mixed corpus annotated with Universal Dependencies POS tags for code-mixed words, and conducts the first comprehensive linguistic analysis of Persian-English code-mixing across multiple social media platforms.

Ghazal Kalhor, Zahra Jafari, Amirarsalan Shahbazi et al. · 0 citations
Open access Aug 2026

Linguistically Informed Machine Learning for Gujarati–English Code-Mixed Sentiment Classification: A Comparative Study of Feature Fusion Strategies

Overall, this work demonstrates that incorporating explicit linguistic information, including language identity, sentiment polarity, and intensifier information, improves sentiment classification of Gujarati–English code-mixed text.

Chirag D. Shah, Shailesh A. Chaudhari · 0 citations
#natural language process... Preprint Sep 2026

TurEngMix: A Text Corpus and Benchmark for Turkish-English Code-Mixed Language Identification and Named Entity Recognition

Natural language processing systems underperform on code-mixed text, particularly for low-resource language pairs. Turkish-English poses a further challenge: it lets English stems combine with Turkish suffixes to form single mixed-language tokens. We introduce TurEngMix, a corpus of 5.5K noisy, naturally occurring soci...

Ilayda Dogan, Phuong-Anh Nguyen-Le, Julia Mendelsohn · 0 citations
Open access Aug 2026

A Distilbert Case-Based Deep Learning Model for Part-of-Speech Tagging for Under-Resourced Kenyan Language: A Case of Dholuo

The results of the experiments reveal that the proposed DistilBERT-based model outperforms the baseline Conditional Random Field and Bidirectional Long Short-Term Memory models with an accuracy 79.05%, precision 80.63%, recall 79.05% and F1-score 79.24%.

Maureen Otieno, L. Wanzare, Calvins Otieno · 0 citations
#natural language process... Preprint Sep 2026

IndicTriMix: Developing Language Identification Datasets and Models for Tri-Language Code-Mixing

This work proposes two approaches of code-mixed generation using parallel sentences of three languages and fine-tune contextual transformer-based models MuRIL and XLM-RoBERTa best suited for Indian languages for language identification in code-mixed settings.

Pruthwik Mishra, Rudra H. Trivedi, Avi Patel et al. · 0 citations
Open access 2026

SwitchEmbed: Representation Learning for Arabic-English Code-Switched Text

This work introduces CS-SNLI, a large-scale Arabic-English code-switched natural language inference dataset and adopts a contrastive learning framework with natural language inference supervision to encourage semantically consistent sentence embeddings across monolingual and code-switched variants.

Mariam Rizkallah, Amani Ghonim, A. Sherif et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.