Vocal Biomarkers for Parkinson’s Disease Detection: A Critical Review of Acoustic Features, Learning Architectures, and the Cross-Dataset Generalisation Problem
Speech impairment affects the large majority of Parkinson’s disease (PD) patients and can appear before motor signs become clinically evident. Two decades of computational research have produced a heterogeneous literature in which reported accuracy varies enormously across studies. This review argues that the spread is governed less by algorithmic capability than by three evaluation choices: subject-level data leakage, leave-one-out cross-validation on very small datasets, and the near-absence of genuinely independent validation. Building on the Ngo et al. 2022 systematic review of the 2010–2021 literature as a verified foundation, this work extends coverage through 2025 with emphasis on self-supervised foundation models and federated learning, both absent from prior surveys. A verified benchmark spanning the validation-rigour spectrum illustrates the effect directly: the foundational sound-booth study of 31 participants reports 91.4% accuracy, whereas the largest telephone-quality study, using a considerably richer feature set and a larger cohort, reports 66.4% balanced accuracy—a 25-point gap attributable to acquisition conditions rather than to method. A structured comparison between handcrafted acoustic features and self-supervised speech embeddings is provided, alongside a three-tier gap analysis linking ten research deficits to prioritised future directions, and a privacy and ethics analysis not present in prior PD voice biomarker surveys. Fine-tuned transformer models achieve the strongest independently validated result to date. The field’s central unsolved problem remains generalisation, not algorithmic sophistication.