V3-Gemma: an on-device multimodal framework for depression screening through clinical-computational alignment
Abstract
Depression is a prevalent mental health disorder that often remains unrecognized in real-world settings, and although artificial intelligence approaches using digital signals show promise for screening, many lack interpretability and rely on cloud-based processing that limits clinical use. This study developed and evaluated V3-Gemma, an on-device multimodal framework for depression screening based on a Clinical-Computational Alignment (CCA) approach that integrates visual, vocal, and verbal cues within a structured clinical reasoning architecture. A total of 130 adults (65 with depression and 65 controls) completed a one-minute picture-description task; 20 observable multimodal features were defined by clinicians and refined to a final set of 19, then implemented as structured prompts for a vision-audio-language model with hierarchical agent-based orchestration, with core inference running locally on-device. In the feature-based analysis, the random forest performed best on an independent test set (50 participants; AUC 0.779, sensitivity 0.84, specificity 0.52). In the criterion-level analysis, features mapped to DSM-5 symptom domains and combined using the DSM-5 rule yielded a test-set accuracy of 0.64 with high sensitivity (0.88) but limited specificity (0.40). Given the small test set, these results represent exploratory feasibility evidence. This proof-of-concept demonstrates a privacy-preserving, interpretable screening approach aligned with DSM-5 symptom domains, pending validation in larger, more diverse samples.