TUNIPROD: An Optimized ProdLDA-Based Multi-Stage Framework for Cross-Lingual Topic Modeling of Code-Switched Public Discourse
Abstract
The rapid proliferation of social data during global crises offers unprecedented opportunities for real-time monitoring; yet, existing research remains constrained by a monolingual bias, often failing in linguistically diverse regions characterized by complex code-switching. This paper addresses this gap by introducing TUNIPROD, an optimized, multi-stage framework designed for cross-lingual topic discovery within the Tunisian digital landscape—a unique environment where Arabic, French, English, and Arabizi (Latin-scripted Arabic) intersect. We present a robust data processing pipeline that aggregates 177,000 documents from diverse sources, including social media, TV, and radio, and employs a multi-stage normalization process specifically tailored for Arabizi. Central to the TUNIPROD framework is the optimization of the Product of Experts Latent Dirichlet Allocation (ProdLDA) model through a systematic grid search of Dirichlet hyperparameters and expert instances. Our results demonstrate that the optimized TUNIPROD framework achieves a high topic coherence score of 0.89, significantly outperforming standard LDA models in capturing the nuanced thematic dependencies of crisis-related discourse. We identify seven dominant macro-topics, revealing that public discourse was primarily driven by Economic (23%) and Health (22%) concerns, with distinct thematic emergence for Vaccination and Governance.