Skip to content
Review Open access

Ordinal sentiment classification in cancer support forums: a controlled benchmark of machine learning and transformer models

Aug 2026 · Frontiers in Digital Health · Vol 8 · 0 citations · 29 references
Medicine

Abstract

Introduction: Online cancer support forums contain naturalistic accounts of fear, uncertainty, coping, and caregiver strain. Such text may contribute to digital phenotyping as one component of longitudinal, human-supervised monitoring, but the operational link between psychosocial distress and sentiment labels requires explicit evaluation. Methods: We benchmarked representative classical, recurrent, and transformer-based models for four-class ordinal sentiment classification using the Mental Health Insights—Vulnerable Cancer Patients dataset (N = 10,392). Models were evaluated under a single validation-guided 60/20/20 holdout split using weighted F1, macro one-vs-rest AUC, class-specific performance, and paired comparisons. Results: Transformer models achieved the strongest overall performance. ALBERT produced the highest weighted F1 and macro AUC (0.7667 and 0.931, respectively), while BioBERT was closely comparable (weighted F1 = 0.7613; macro AUC = 0.917) and showed slightly higher recall for the “very negative” class (0.8019 vs. 0.7736). Error analysis showed that transformer errors concentrated around ordinal decision boundaries, while residual positive-class errors remained operationally important for supportive workflows. Discussion: These split-specific findings support transformer fine-tuning as a decision-support component for vulnerability-oriented monitoring, while emphasizing calibration, transparent error review, and human oversight rather than autonomous clinical assessment.

Read PDF

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.