Vocal Biomarkers for Parkinson’s Disease Detection: A Critical Review of Acoustic Features, Learning Architectures, and the Cross-Dataset Generalisation Problem
Abstract
Speech impairment affects the large majority of Parkinson’s disease (PD) patients and can appear before motor signs become clinically evident. Two decades of computational research have produced a heterogeneous literature in which reported accuracy varies enormously across studies. This review argues that the spread is governed less by algorithmic capability than by three evaluation choices: subject-level data leakage, leave-one-out cross-validation on very small datasets, and the near-absence of genuinely independent validation. Building on the Ngo et al. 2022 systematic review of the 2010–2021 literature as a verified foundation, this work extends coverage through 2025 with emphasis on self-supervised foundation models and federated learning, both absent from prior surveys. A verified benchmark spanning the validation-rigour spectrum illustrates the effect directly: the foundational sound-booth study of 31 participants reports 91.4% accuracy, whereas the largest telephone-quality study, using a considerably richer feature set and a larger cohort, reports 66.4% balanced accuracy—a 25-point gap attributable to acquisition conditions rather than to method. A structured comparison between handcrafted acoustic features and self-supervised speech embeddings is provided, alongside a three-tier gap analysis linking ten research deficits to prioritised future directions, and a privacy and ethics analysis not present in prior PD voice biomarker surveys. Fine-tuned transformer models achieve the strongest independently validated result to date. The field’s central unsolved problem remains generalisation, not algorithmic sophistication.