Skip to content
Open access

V3-Gemma: an on-device multimodal framework for depression screening through clinical-computational alignment

Aug 2026 · Scientific Reports · 0 citations

Abstract

Depression is a prevalent mental health disorder that often remains unrecognized in real-world settings, and although artificial intelligence approaches using digital signals show promise for screening, many lack interpretability and rely on cloud-based processing that limits clinical use. This study developed and evaluated V3-Gemma, an on-device multimodal framework for depression screening based on a Clinical-Computational Alignment (CCA) approach that integrates visual, vocal, and verbal cues within a structured clinical reasoning architecture. A total of 130 adults (65 with depression and 65 controls) completed a one-minute picture-description task; 20 observable multimodal features were defined by clinicians and refined to a final set of 19, then implemented as structured prompts for a vision-audio-language model with hierarchical agent-based orchestration, with core inference running locally on-device. In the feature-based analysis, the random forest performed best on an independent test set (50 participants; AUC 0.779, sensitivity 0.84, specificity 0.52). In the criterion-level analysis, features mapped to DSM-5 symptom domains and combined using the DSM-5 rule yielded a test-set accuracy of 0.64 with high sensitivity (0.88) but limited specificity (0.40). Given the small test set, these results represent exploratory feasibility evidence. This proof-of-concept demonstrates a privacy-preserving, interpretable screening approach aligned with DSM-5 symptom domains, pending validation in larger, more diverse samples.

Read PDF

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.