Skip to content
Open access

The use of large language models in automated depression detection.

Aug 2026 · Acta Psychologica · Vol 269, pp. 107602 · 0 citations · 43 references
Medicine

Abstract

Background

Large language models have been evaluated on many healthcare tasks, including depression screening. However, it is unclear whether estimates of performance are accurate, especially in a setting with realistic clinical constraints.

Methods

We use the publicly-available Distress Analysis Interview Corpus - Wizard of Oz (DAIC-WOZ) to give performance estimates of LLMs that respect patient privacy and could be feasibly deployed in a clinical setting. These models are locally run and under 15 billion parameters.

Results

Accuracy, sensitivity, and specificity of the models we evaluated ranged from 0.233-0.677, 0.041-0.929, 0.000-0.729, respectively. There are significant differences in the performance we observed versus other studies that evaluate commercial models. We also demonstrate poor agreement amongst different LLMs.

Conclusion

Current performance estimates of LLMs with respect to depression screening are most likely optimistic. When restricted to smaller models that could be locally deployed (for privacy protection) in a clinical setting, LLMs do not detect depression with sufficient accuracy, sensitivity, or specificity to be used in a screening programme.

Read PDF