Aug 2026· Frontiers in Big Data· Vol 9· 0 citations· 18 references
Medicine
TL;DR
The first structured analysis of micro macro divergence and rare class behavior in English Malayalam code mixed POS tagging is provided, providing the first structured analysis of micro macro divergence and rare class behavior in English Malayalam code mixed POS tagging.
Abstract
Introduction Part Of Speech (POS) tagging is a fundamental task in Natural Language Processing (NLP) that assigns grammatical labels to words in a sentence. Code mixed text, which entails switching between two or more languages within a single conversation or a sentence, presents challenges for POS tagging. This investigation entailed a comprehensive study of deep learning approaches for cross linguistic POS tagging, focused on English Malayalam code mixed data prevalent on social media platforms. The study was carried out on linguistically complex and varied English Malayalam code mixed text from social media platforms with informal spellings, language switching, transliteration, slang, and unclear grammatical boundaries, reflecting the characteristics of informal online communication. We provide the first structured analysis of micro macro divergence and rare class behavior in English Malayalam code mixed POS tagging. Methods We evaluated 14 state of the art model configurations that span traditional sequence labeling approaches and multilingual transformer architectures. Models were compared using standard performance metrics prevalent in the domain of data science, supplemented by normalized confusion matrices, error prone tag identification and micro-macro F1 gap analysis. Results Our results showed that CRF (No Lang) emerged as the most balanced model overall on macro F1 (all classes) of 0.8170, while (BiLSTM + CRF) achieved the highest macro F1 (seen classes) of 0.8831, precision of 0.9167, and recall of 0.875, though this reflects strong performance concentrated on frequent tag classes rather than balanced coverage across the full tag set. Notably, the pretrained multilingual transformers (mBERT, MuRIL), despite prior exposure to Malayalam during pretraining, were outperformed on several key metrics by CRF and BiLSTM models trained directly on the code mixed dataset. Discussion This finding was contrary to our expectation that existing multilingual knowledge would translate into a clear advantage on this task and merits further investigation.
This work introduces PERCEPT, the first publicly available large-scale Persian-English code-mixed corpus annotated with Universal Dependencies POS tags for code-mixed words, and conducts the first comprehensive linguistic analysis of Persian-English code-mixing across multiple social media platforms.
Ghazal Kalhor, Zahra Jafari, Amirarsalan Shahbazi et al.· 0 citations
Overall, this work demonstrates that incorporating explicit linguistic information, including language identity, sentiment polarity, and intensifier information, improves sentiment classification of Gujarati–English code-mixed text.
Chirag D. Shah, Shailesh A. Chaudhari· International journal of com...· 0 citations
Natural language processing systems underperform on code-mixed text, particularly for low-resource language pairs. Turkish-English poses a further challenge: it lets English stems combine with Turkish suffixes to form single mixed-language tokens. We introduce TurEngMix, a corpus of 5.5K noisy, naturally occurring soci...
Ilayda Dogan, Phuong-Anh Nguyen-Le, Julia Mendelsohn· 0 citations
The results of the experiments reveal that the proposed DistilBERT-based model outperforms the baseline Conditional Random Field and Bidirectional Long Short-Term Memory models with an accuracy 79.05%, precision 80.63%, recall 79.05% and F1-score 79.24%.
Maureen Otieno, L. Wanzare, Calvins Otieno· International Journal of Com...· 0 citations
This work proposes two approaches of code-mixed generation using parallel sentences of three languages and fine-tune contextual transformer-based models MuRIL and XLM-RoBERTa best suited for Indian languages for language identification in code-mixed settings.
Pruthwik Mishra, Rudra H. Trivedi, Avi Patel et al.· 0 citations
This work introduces CS-SNLI, a large-scale Arabic-English code-switched natural language inference dataset and adopts a contrastive learning framework with natural language inference supervision to encourage semantically consistent sentence embeddings across monolingual and code-switched variants.
Mariam Rizkallah, Amani Ghonim, A. Sherif et al.· International Conference on...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.