The findings demonstrate that transformer-based multi-label learning can support scalable, reproducible analysis of HIV-related health perceptions in online communities, with potential applications in public health surveillance, communication strategy design, and digital intervention planning.
Abstract
This paper presents a supervised multi-label framework for detecting multidimensional perceived risk in HIV-related Reddit discourse. A longitudinal corpus of 329,707 texts collected from r/hivaids and r/HIV between 2015 and 2025 was analyzed to identify three risk dimensions: transmission risk, health deterioration risk, and social stigma risk. A stratified sample of 2,000 texts was annotated by domain experts, achieving substantial inter-annotator agreement (Cohen's κ = 0.74-0.81). A RoBERTa-base model was fine-tuned using class-weighted binary cross-entropy loss and per-class threshold optimization. The proposed model achieved a macro-F1 score of 0.87 and a macro-AUC-ROC of 0.97, outperforming 12 baseline models, including traditional machine learning, neural network, and alternative transformer-based approaches. Ablation experiments confirmed the importance of transformer fine-tuning and class weighting, while also showing that handcrafted features provided only marginal gains. Applied to the full corpus, the model revealed significant upward trends in transmission risk and health deterioration risk, strong co-occurrence between transmission and stigma-related discourse, and distinct information-seeking patterns across risk categories. The findings demonstrate that transformer-based multi-label learning can support scalable, reproducible analysis of HIV-related health perceptions in online communities, with potential applications in public health surveillance, communication strategy design, and digital intervention planning.
Mental health risk detection from user-generated social media text has become increasingly important as psychiatric conditions such as depression, anxiety, and suicidal ideation continue to rise. However, most existing computational studies operationalize the problem using single-label datasets, implicitly assuming mutually exclusive conditions and thereby under-modeling clinically prevalent comorbidity. To better align modeling assumptions with real-world mental health phenomena, we formulate social-media-based risk identification as a multi-label text classification problem, enabling the simultaneous prediction of co-existing risks within a single post. We construct a hybrid corpus by integrating AIMH/SWMH and the Sentiment Analysis for Mental Health dataset, and by adding neutral/positive samples from Sentiment140 to strengthen healthy-content representation. Multi-label annotations are generated via Meta Llama-3-70B-Instruct using deterministic zero-shot prompting (temperature=0); a manual audit of 1,000 randomly sampled instances yields 94% agreement with the LLM-generated labels. We then benchmark multiple transformer-based encoders under a unified multi-label training protocol and report macro/micro F1 performance, highlighting the feasibility of transformer-based multi-label learning for modeling overlapping mental health risks at scale in social media.
Unknown authors· El-Cezeri Fen ve Mühendislik...· 0 citations
Improvements in recall and F1-score for minority classes demonstrate the effectiveness of the balancing process and highlight the potential of machine learning for student mental health classification, although further validation on larger and more diverse datasets is required.
It is shown that con-textual transformer/LLM models yield more reliable macro-level performance under imbalance than TF–IDF baselines, particularly for semantically adja- cent classes.
Nehal Shah, Mehul Barot· International journal of com...· 0 citations
Introduction: Online cancer support forums contain naturalistic accounts of fear, uncertainty, coping, and caregiver strain. Such text may contribute to digital phenotyping as one component of longitudinal, human-supervised monitoring, but the operational link between psychosocial distress and sentiment labels requires explicit evaluation. Methods: We benchmarked representative classical, recurrent, and transformer-based models for four-class ordinal sentiment classification using the Mental Health Insights—Vulnerable Cancer Patients dataset (N = 10,392). Models were evaluated under a single validation-guided 60/20/20 holdout split using weighted F1, macro one-vs-rest AUC, class-specific performance, and paired comparisons. Results: Transformer models achieved the strongest overall performance. ALBERT produced the highest weighted F1 and macro AUC (0.7667 and 0.931, respectively), while BioBERT was closely comparable (weighted F1 = 0.7613; macro AUC = 0.917) and showed slightly higher recall for the “very negative” class (0.8019 vs. 0.7736). Error analysis showed that transformer errors concentrated around ordinal decision boundaries, while residual positive-class errors remained operationally important for supportive workflows. Discussion: These split-specific findings support transformer fine-tuning as a decision-support component for vulnerability-oriented monitoring, while emphasizing calibration, transparent error review, and human oversight rather than autonomous clinical assessment.
Zhongyan Wang, Yuchen Cao, Shuo Xu et al.· Frontiers in Digital Health· 0 citations
Social-media posts record not only what users say but also when and how they participate. These two sources of evidence can support computational screening for depression-related patterns, although the task is complicated by indirect language, overlapping class boundaries, and incomplete behavioral records. This study examines whether combining information at several levels produces more informative user representations than relying on isolated posts. The proposed workflow begins with BERT-based encoding of post content. Random downsampling is used in the pre-trained-model experiments to reduce the imbalance between the two labels. Posting frequency, activity-time distribution, and historical-post volume are then mapped into the same feature space as the text representation. Post-level vectors are aggregated for each user, and a graph neural network is applied to a similarity graph so that the final representation also reflects relationships among users. A supervised contrastive objective is included to encourage compact within-class representations and clearer separation between classes. Across the reported test results, user-level models were markedly stronger than tweet-level baselines. The user-level BERT + GNN configuration obtained an accuracy of 0.9404, an F1-score of 0.9399, and an AUC of 0.9703. These results suggest that historical context and relational structure are useful for this classification setting; they should not, however, be interpreted as a substitute for clinical assessment.
Xu Zhang· Frontiers in Computing and I...· 0 citations
Existing university mental health monitoring often depends on voluntary help-seeking or manual questionnaire interpretation, which may delay early support for students experiencing academic stress. This study proposes an explainable XGBoost-based early-warning framework for non-clinical mapping of student mental health risk from academic stress indicators. The single-site dataset comprised 1,002 anonymized student records from Universitas Muria Kudus. K-Means clustering was used to transform DASS-21 depression, anxiety, and stress scores into low, moderate-, and high-risk categories, while XGBoost predicted the cluster-derived labels using seven single-item academic stress indicators and engineered aggregate and interaction features. On a stratified hold-out testing set of 201 records, the model achieved weighted precision, recall, and F1-score values of 0.8907, 0.8905, and 0.8906, respectively, with class-level F1-scores of 0.9109 for low risk, 0.8900 for moderate risk, and 0.8713 for high risk. Additional ablation, clustering sensitivity, subgroup, threshold, and SHAP stability analyses were conducted to strengthen robustness and interpretability. The findings show that cumulative academic stress and interaction features involving parental expectations, exam anxiety, and learning-method adaptation were consistently influential predictors. The framework is intended to support early institutional prioritization and counseling referral, not clinical diagnosis. Generalization remains limited by the single-institution sample and the use of single-item academic stress indicators; therefore, local retraining and recalibration are required before institutional deployment, including implementation of the Streamlit prototype.
Supriyono, H. Firmansyah, Soni Adiyono· Jurnal RESTI (Rekayasa Siste...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.