Ensemble Learning for Multi-Source Multi-Domain Sentiment Analysis
Abstract
Sentiment analysis systems are typically trained and evaluated on a single data source and domain, limiting their reliability when applied across sources that differ in vocabulary, review length, and writing style. This paper presents an ensemble learning framework for multi-source, multi-domain sentiment classification evaluated across three structurally distinct sources: Twitter (short, informal), TripAdvisor (long, first-person), and Rotten Tomatoes (short, editorial). A unified pipeline combining text normalization with a TF-IDF-weighted Word2Vec (Skip-gram, 200-dim) representation was used to train six base classifiers (Naive Bayes, KNN, Logistic Regression, SVM, Decision Tree, MLP) and four ensembles (Bagging, Boosting, Stacking, Voting) under an identical 5-fold cross-validation protocol. On a held-out, domain-stratified test set of 2,360 samples, SVM achieved the best overall performance (83.05% accuracy, 82.28% F1), narrowly outperforming Stacking (82.58%) and Voting (81.78%). A domain-wise breakdown revealed that TripAdvisor (88.13% F1) was classified far more reliably than Rotten Tomatoes (78.41%) or Twitter (76.27%), a gap associated with review length rather than dataset size. These findings show that ensembling does not guarantee improvement over an already-strong individual classifier, and that domain-aware evaluation is essential in multi-source sentiment analysis.