Skip to content
Open access

Evaluating CNN Performance for Sentiment Classification of Mental Health-Related Social Media Text: A Comparative Study of Small and Large Datasets

2026 · International journal of research and scientific innovation · 0 citations

Abstract

Social media provides a large amount of user-generated text in which people share their thoughts, feelings, experiences, and opinions. This text can be studied using natural language processing and deep learning techniques. This study evaluates the performance of a Convolutional Neural Network (CNN) for sentiment classification of mental-health-related social media text. A small dataset and a larger dataset are compared to examine the classification performance of the CNN across datasets of different sizes. The small dataset contains 732 social media text entries, while the large dataset contains 27,972 entries. The text data were prepared using lowercasing, tokenization, stopword removal, stemming, and lemmatization. TF-IDF, GloVe, and Word2Vec were used for text representation as described for the respective datasets. CNN models were developed using Keras and evaluated using accuracy, precision, recall, and F1-score. The CNN achieved an accuracy of 97% on the small dataset and 90.56% on the large dataset. The reported F1-scores were 49% and 90%, respectively. The difference between accuracy and F1-score in the small dataset shows the importance of considering class distribution when evaluating classification performance. The results indicate that CNN can classify sentiment patterns in mental-health-related social media text. However, sentiment classification cannot be considered a clinical diagnosis or direct evidence of schizophrenia. The study therefore focuses on sentiment classification and on comparing CNN performance across datasets of different sizes.

Read PDF

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.