Large Language and Ensemble Powered Enhanced Parkinson's Detection Model
Abstract
Parkinson’s disease (PD) is a progressive neurodegenerative disorder that severely impairs motor control and quality of life. Conventional diagnostic approaches such as clinical motor assessment and neuroimaging are expensive, invasive, and lack sensitivity for early-stage detection. However, subtle alterations in vocal characteristics, including tremors, pitch instability, and hoarseness, can manifest years before motor symptoms become evident, offering a promising non-invasive biomarker. To address these challenges, we propose a novel ensemble-driven machine learning framework that leverages acoustic features from speech signals for early PD detection. Using an open-access voice dataset, 22 clinically validated vocal biomarkers were extracted. Six supervised models were trained and evaluated, with a custom Stacking Classifier achieving superior performance with an accuracy of 97.77%, precision of 98.39%, recall of 95.31%, and an F1-score of 96.83%. This ensemble method integrates the strengths of diverse learners to ensure stable and generalizable predictions. Furthermore, the system incorporates a post-diagnostic chatbot powered by a Large Language Model (LLM), enabling personalised support and guidance. The proposed framework is non-invasive, scalable, and well-suited for telemedicine, representing a significant step toward accessible early-stage PD screening.