The system reduces supervisory triage latency from 72 hours to real time (~10 seconds per session), enabling proactive intervention in high-risk cases and addresses the cold-start problem through Bayesian priors and implements timestamp-based modality synchronization for robust multi-modal fusion.
Abstract
Modern mental healthcare faces a critical shortage of senior supervisory oversight, leading to a"supervision gap"where novice therapists manage high-stakes risks with delayed professional feedback. This paper proposes a new framework utilizing a fine-tuned Mistral-7B-instruct model as an automated"Supervisor-in-the-Loop"system. By leveraging 106 sessions from the DAIC-WOZ dataset, the model performs a tri-stream analysis: (1) Therapeutic Alliance tracking via semantic adherence, (2) Latent risk prediction using attention-weighted analytics, and (3) Supervisory Triage via a Dynamic Clinical Urgency Index (D-CUI). Our multi-modal VAL (Visual-Acoustic-Linguistic) framework achieves 95% technique identification accuracy [95% CI: 75.1%-99.9%], alliance assessment MAE of 0.105 on a 5-point scale [95% CI: 0.059-0.151], therapeutic fidelity alpha = 0.423, and mean D-CUI of 0.370 [95% CI: 0.322-0.419]. Training converged in 105 steps with 85.2% loss reduction on a single Tesla T4 GPU. The system reduces supervisory triage latency from 72 hours to real time (~10 seconds per session), enabling proactive intervention in high-risk cases. The system addresses the cold-start problem through Bayesian priors and implements timestamp-based modality synchronization for robust multi-modal fusion.
This paper presents a system for the CLPsych 2026 Shared Task on longitudinal mental health modeling from social media timelines, grounded in the MIND framework (Atzil-Slonim, 2025). MIND conceptualizes mental health as evolving self-states defined by A ffect, B ehavior, C ognition, and D esire (ABCD), providing a structured lens on mental health trajectories. The system centers on a theory-explicit prompting framework for structured sequence summarization (Task 3.1) and recurrent dynamic signature extraction (Task 3.2), encoding the full ABCD taxonomy directly into the LLM prompt to ensure clinically grounded, inter-pretable outputs. A three-stage pipeline infers a direction-of-change label per sequence, produces structured ABCD summaries with few-shot exemplar augmentation, and aggregates these summaries to derive cross-individual recurrent patterns. The system ranks first on deterioration-related recurrent signatures and second overall, achieving the top Fit and Specificity scores in Task 3.2, demonstrating the benefits of explicit clinical grounding for conceptual accuracy.
Pawan Kumar, Ankit Meshram, S. Jha et al.· Workshop on Computational Li...· 1 citation
Personalized interpretation of medical reports has emerged as an increasingly important need among patients. Addressing this need requires both evidence-grounded medical factuality and context-dependent patient communication, yet existing medical vision-language tasks do not adequately capture these dual requirements. To bridge this gap, we introduce Patient-oriented Medical Report Interpretation (PMRI), a novel open-ended multimodal generation task that requires models to explain medical reports in accurate and accessible language based on a user's query and dialogue history. These two objectives differ fundamentally in their verifiability, yet remain tightly coupled, making them difficult to optimize jointly under conventional supervised fine-tuning and holistic reinforcement learning paradigms. To address this challenge, we propose G-CARL, a grounded, checklist-aligned reinforcement learning framework that combines multi-source retrieval for atomic claim verification with context-aware, instance-specific weighted checklists for response coverage, providing structured supervision for factuality, user-demand satisfaction, and expression quality without constraining response diversity. We further construct MMedReport, a real-world PMRI benchmark, along with a clinician-designed three-dimensional evaluation protocol. Extensive experiments demonstrate that G-CARL consistently outperforms existing post-training baselines in overall quality, claim-level precision, and checklist recall. Pairwise preference evaluation by clinicians further confirms that G-CARL produces interpretations that are more accurate and better aligned with patient needs.
Shiao Xie, Siyu Chen, Jianwei Lv et al.· 0 citations
Safety-critical mental-health support systems must distinguish when supportive conversation is appropriate from when free-form generation should be blocked. This paper presents Anian, a safety-gated multimodal AI backend for perinatal mental-health support and mindfulness-intervention routing. Anian is not intended to diagnose psychiatric conditions or replace clinical care or crisis intervention. Its modular pipeline places generative AI downstream of structured state representation, conservative risk fusion, and response gating. User text or voice-derived ASR transcripts are mapped into four linked layers: L1 emotion states, L2 psychosocial constructs, L3 safety risk, and L4 intervention routes. Local text- and rule-based safety evidence is fused with external voice-derived evidence using a highest-risk-priority rule, S_fusion = max(S_local, S_external). At moderate or high fused risk, ordinary AI-generated responses and text-to-speech delivery are blocked and replaced by fixed safety content and prompts for human support. An internal prototype evaluation used approximately 858,295 normalized records from public emotion, dialogue, mental-health-related, and Chinese dialogue corpora within a weak-label and rule-derived framework. Micro-F1 scores were 0.9604 for L1 emotion classification, 0.9144 for L2 psychosocial constructs, and 0.9742 for L4 routing. In a controlled safety stress test of 233 samples, the L3 rule engine achieved high-risk recall of 1.0000 within predefined scenarios. These findings support the internal feasibility of the label framework and gating logic but do not establish clinical validity, diagnostic accuracy, real-world safety, or effectiveness. We report the architecture, ontology, safety-fusion mechanism, prototype evaluation, error-analysis plan, and roadmap for expert-reviewed and real-world validation.
Mental health disorders represent a growing global public health concern, yet early detection remains challenging due to the reliance on conventional clinical assessments. The rise of social media provides a unique opportunity to monitor individuals' emotional states through their textual expressions passively. However, existing automated detection approaches are predominantly opaque "black-box" systems, limiting their adoption in clinical and educational settings where interpretability is paramount. This paper presents a comprehensive, multi-modal pipeline for automated mental health risk detection that integrates contextual embeddings from BERT (Bidirectional Encoder Representations from Transformers) with behavioral and linguistic indicators. We benchmark five distinct classifier architectures, including TF-IDF-based and BERT-based models, using Stratified 5-Fold Cross-Validation across multiple metrics: Accuracy, Macro F1-Score, Cohen's Kappa (κ), and Matthews Correlation Coefficient (MCC). The best-performing model, BERT + Logistic Regression, achieves a Macro F1-Score of 0.6464 ± 0.0119 and a Macro AUC-ROC of 0.8726, while BERT + Random Forest achieves the highest accuracy of 0.8742 ± 0.0020. To address the interpretability gap, we implement a multi-faceted Explainable AI (XAI) framework comprising Shapley Additive exPlanations (SHAP) for global and local feature attribution, Local Interpretable Model-agnostic Explanations (LIME) for word-level insights, and Counterfactual Analysis for decision-boundary understanding. Statistical significance is rigorously validated via McNamara’s test(p <0.001) and the Wilcoxon signed-rank test. Our results demonstrate that the integration of rich contextual embeddings with transparent XAI produces a trustworthy, accurate, and actionable system for mental health risk monitoring.
Divya N., V. J. Chakravarthy, K. Jayabharathi et al.· International journal of com...· 0 citations
This Attention Deficit Hyperactivity Disorder (ADHD) remains substantially under-diagnosed among university students despite affecting 2–8% of this population. Campus health services, facing persistent resource constraints, frequently accumulate assessment backlogs of 6–12 months. This paper presents a machine learning framework for automated ADHD pre-screening that combines structured psychometric assessments with natural language processing (NLP)-derived features extracted from free-text clinical self-reports. Drawing on 506 university student responses, we engineer 124 multimodal features spanning four validated instruments, the Adult ADHD Self-Report Scale (ASRS), Beck Anxiety Inventory (BAI), Beck Depression Inventory (BDI-II), and Adult Attachment Scale (AAS), together with unstructured diagnostic text. Mutual Information-based feature selection reduces dimensionality to 20 features, yielding a 2% accuracy gain. A comparative evaluation across five classifiers reveals Logistic Regression as the top performer, achieving 81.4% accuracy and an AUC of 0.881. SHAP (SHapley Additive exPlanations) analysis confirms clinical meaningfulness by identifying BAI Item 8 (somatic anxiety), ASRS inattention items, and prior mental health history as the principal risk factors. The system is deployed as an interactive web application that delivers calibrated risk assessments suited to clinical triage in resource limited settings.
: Artificial Intelligence (AI) is increasingly deployed across the mental health pathway, from screening and diagnosis through to intervention, monitoring and prognosis. This paper presents a structured review, aligned with the Preferred Reporting Items for Systematic Reviews and Meta-Analyses (PRISMA) 2020 guidance, of the current evidence base for AI in mental health. Working from four anchor reviews and a transparent, criteria-driven corpus of supporting primary and regulatory sources, we make four contributions. First, we propose a faceted taxonomy that classifies any mental-health AI system across paradigm, data modality, clinical task, autonomy/risk and evidence maturity. Second, we report an explicit search protocol with inclusion and exclusion criteria, so that the evidence base is reproducible rather than implicit. Third, we synthesise reported performance comparatively across application domains and grade the maturity of the evidence using a five-level scheme. Fourth, we propose a Responsible AI Pipeline that connects research activity to safe clinical deployment through explicit bias, validation, safety and regulatory gates. Reported strengths include accurate classification and risk prediction for common mental disorders, earlier case-finding, scalable chatbot-based self-help, and support for personalised treatment planning. However, the literature remains marked by methodological inconsistency, limited external validation, bias in training data, under-representation of people with intellectual disability and other marginalised groups, and unresolved issues around consent, explainability and regulation. We argue that AI should be framed as an augmentation of - not a replacement for - the clinical relationship, with equity, consent and explainability treated as first-order design constraints.
Dr. B. Anand, Poornima Ramachandran, Abhishek Subramaniam· Proceedings of the Internati...· 0 citations