Skip to content
Open access

68. Predicting executive dysfunction in mood disorders using a multimodal machine learning approach

Sep 2026 · International Journal of Neuropsychopharmacology · Vol 29, pp. i99 - i99 · 0 citations

Abstract

Abstract Background Mood disorders are often accompanied by persistent cognitive impairment, even during remission. Recent advances in computational psychiatry may serve as digital biomarkers of cognition. Aims & Objectives This study aimed to develop machine learning models using multimodal features derived from patients’ spontaneous speech, to predict executive function. Method This study recruited 196 participants, including patients with unipolar depression, bipolar disorder, and healthy controls, from hospital in northern Taiwan . Each participant completed the Wisconsin Card sorting Test (WCST) to assess cognitive function. We focused on the percentage of total errors (PTE) and labeled as WCST-PTE score. The WCST-PTE score underwent standardization and the cut-off point is 40. Participants were prompted with an open question “Please describe your recent emotions, life circumstances, and mental states” and their spontaneous speech was recorded during euthymic state. Two types of audio features were extracted:(1) Handcrafted features (2) Learned features. For text features, both handcrafted and learned embeddings were used. Demographic variables were also included. Multiple machine learning classifiers were trained to predict cognitive status. A 5-fold cross-validation repeated 10 times was employed to ensure robustness, and random oversampling was applied to balance class distribution in the training folds. Results In the WCST-PTE analysis, 196 participants were recruited ( normal : 155 (79.1%) , impaired: 41 (20.9%)). Among the impaired group, 85.4% were mood disorder patients. Baseline demographics showed that WCST-PTE performance was significantly associated with multiple variables: lower education level, shorter years of education, older age, psychiatric hospitalization history (p<0.01) . Demographic(Dem)-only models using XGB achieved a baseline performance of AUC = 0.683 for predicting cognition impairment. Adding audio features significantly enhanced prediction. The prosodic + demographic model reached AUC = 0.715 (p < 0.001), while the linguistic inquiry and word count (LIWC) + demographic text-based model showed modest improvement (AUC = 0.709, p < 0.001). A multimodal model combining prosodic, linguistic, and demographic features (PF + LIWC + Dem) yielded AUC = 0.713, but did not significantly outperform PF + Dem (p = 0.504). Discussion & Conclusions The WCST-PTE findings demonstrate that acoustic-prosodic markers, combined with simple demographic information, may reliably detect executive dysfunction.

Read PDF

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.