Author

G. Kancharla

1 paper indexed here

Fetches their full publication history.

Not the right person? Other researchers publish under this name.

Open access Jun 2026

Language Identification in Transliteration-Based Code-Mixed Text: A Study on Telugu–English Data

In recent years, social media users in multilingual regions have begun mixing languages more freely, and Telugu–English combinations are among the most common examples in India. Much of this content appears in informal Roman transliteration, and the lack of uniform spelling makes automatic processing difficult. In this work, we focus on word-level Language Identification (LID) for such transliterated text. Our approach relies on character-based TF–IDF features and a set of traditional machine-learning models. In this study, we worked with four different models—Multinomial Naive Bayes, Logistic Regression, Random Forest, and Support Vector Machine—and evaluated them on an annotated set that included Telugu, English, Named Entity, and Universal tokens. Among the four, the SVM turned out to be the strongest, reaching an accuracy of 86% along with an F1-score of 0.85. The study also brings out some practical issues with real-world transliterated text, particularly class imbalance and the wide range of spelling variations. We conclude with possible directions for improvement, including the use of neural and transformer-based models that might capture more contextual cues in future versions of this system.

Adarshavathi Jampala, Padmavathi Guddetti, G. Kancharla · 0 citations