Longitudinal Artificial Intelligence for Parkinson’s Disease: Modern Modelling, Feature Selection, and Imputation – A Technical Review and Clinical Roadmap - Supplementary Material
Abstract
Parkinson’s disease (PD) research has been reshaped by large-scale multi-modal longitudinal datasets and advances in artificial intelligence (AI) and machine learning (ML). This review synthesises the current state of longitudinal modelling, feature selection, and missing-data imputation for PD, drawing on 154 studies identified through a documented search strategy covering January 2011 to July 2025, with a focus on technical rigour and clinical translation. Classical and modern approaches for modelling disease progression are critically examined, including linear and latent-class mixed-effects models, latent-growth and event-based frameworks, and ML methods such as Subtype and Stage Inference (SuStaIn), graph-based neural networks, and neural ordinary differential equations. Feature-selection paradigms are discussed from statistical and AI-driven perspectives, with attention to cost-aware, fairness-aware, and multi-omics strategies. Imputation methods, from simple techniques to deep generative models, are then reviewed, together with an analysis of how imputation choices cascade through downstream modelling, feature selection, and clinical inference. Reported developments include latent-class mixed-effects models reaching an of 0.92 for cognitive-decline prediction in single-cohort Parkinson’s Progression Markers Initiative (PPMI) analyses, multimodal SuStaIn-based subtyping refined with deep learning and multi-omics integration, graph-based personalised progression models achieving area under the receiver operating characteristic curve (AUCs) of 0.71–0.74 on PPMI/PDBP test sets, and multi-objective feature-selection frameworks jointly optimising cost, equity, and interpretability. Recurring gaps include limited external validation across diverse populations, inconsistent leakage prevention, and weak characterisation of how imputation choices propagate to downstream inference. Priority directions include causal-probabilistic models, uncertainty-driven feature acquisition, federated validation, and transparent imputation protocols with routine fairness auditing, aligned with current reporting standards for systematic reviews and clinical prediction models.