Text sentiment analysis of the social media text faces challenges posed by unstructured data and labori- ous human labeling for intent-driven, hierarchical classification. This work compares conventional ML models (SVM, Naïve Bayes, Logistic Regression) with contextual DL models (DistilBERT) in terms of their performance on Sentiment140 dataset (1.6 million tweets) where a balanced 300,000 tweets were selected (200,000 training, 50,000 validation and 50,000 test). As a solution to the bottleneck of human labeling for detailed topic classification, a zero-shot classification pipeline that uses Natural Language Inference (NLI) for classification of 1,000 tweets into a two-layer taxonomy of 30 parent topics and 330 subtopics has been created without any human-labeled samples. SVM is able to achieve 81.52% accuracy, while DistilBERT scores 84.44% on 50,000 tweet test set and 83.0% on a small 1,000 tweets sample.
Bhumit Peshavariya, S. Nahar· Informatica· 0 citations
Considering the growth of multilingual user made content within social-media platforms, there is an urgent need for developing scalable, language-agnostic approaches for their analysis. Within this paper, we analyze mBERT's performance in sentiment classification in a binary setting as well as the possibility of performing transfer learning between languages. Specifically, the fine-tuned model is applied for sentiment analysis of tweets from the preprocessed TweetEval dataset, obtaining 79.2% of accuracy and 74.7% of the F1 score. It is shown that cross-language transfer learning without any preliminary training on multilingual sentiment datasets provides quite satisfactory performance. However, a more complex approach can be used, which consists of applying filtering of negative sentiments, categorization of subcategories through a sentence transformer with zero-shot settings, and grouping the resulting data in several major categories to obtain severity scores according to frequency thresholds. The application of the sentiment classification with transformers in combination with issue prioritization makes it possible to develop an end-to-end approach to structuring multilingual social media content.
S. Nahar, P. P. Agnihotri· International Journal of Sci...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.