Machine learning methods for predictive modeling in personalised medicine
TL;DR
This dissertation demonstrates how biologically informed machine learning and representation learning approaches can support scalable predictive modeling and transcriptomics-driven therapeutic inference across diverse biomedical applications.
Abstract
The increasing availability of large-scale molecular, physiological, and clinical data has transformed biomedical research and accelerated the development of personalised medicine. At the same time, the complexity and heterogeneity of these data introduce major computational and methodological challenges, including high dimensionality, noise, limited interpretability, and difficulties in integrating information across diverse biological modalities. Therefore, machine learning has emerged as a central framework for extracting clinically meaningful information from complex biomedical systems and for supporting predictive and therapeutic decision-making. The objective of this dissertation is to develop machine learning methods for predictive and therapeutic modeling in personalised medicine, enabling the extraction of clinically relevant insights from heterogeneous biomedical data. In particular, this work focuses on two complementary directions: predictive modeling from physiological signals and transcriptomics-driven therapeutic modeling using biologically structured representation learning approaches. For physiological signal analysis, this dissertation investigates the development of computationally efficient neural network architectures for atrial fibrillation prediction from electrocardiography signals. In contrast to approaches that primarily optimize predictive performance, the proposed framework integrates deep learning with hardware-aware optimization in order to enable deployment within resource-constrained environments such as embedded biomedical systems. By jointly considering predictive accuracy, computational complexity, and energy efficiency, the resulting models support scalable real-time physiological monitoring while maintaining robust predictive performance. A central focus of this dissertation is the development of machine learning frameworks for therapy prediction directly from transcriptomic data. To address this problem, a hierarchical representation learning framework was developed to infer drug mechanisms of action from gene expression perturbation signatures. Using supervised contrastive learning, the model organizes transcriptional responses into biologically structured latent spaces that capture both high-level mechanistic organization and compound-specific transcriptional substructure. This representation learning strategy improves interpretability and generalization across unseen compounds and cellular contexts while preserving biologically meaningful transcriptional relationships. Building upon this latent-space formulation, we introduce RANKOR, a machine learning framework for direct drug prioritization from bulk and single-cell transcriptomic signatures. Unlike traditional enrichment-based methods that depend on predefined perturbational reference signatures, RANKOR learns aligned transcriptomic and chemical latent spaces that enable scalable therapeutic ranking directly from molecular signatures. By integrating transcriptomic and chemical representations within a shared latent framework, the method enables prioritization of transcriptionally unseen compounds directly from their chemical structure while maintaining competitive predictive performance and substantially reduced computational cost relative to classical enrichment-based approaches. In addition to these methodological contributions, the dissertation includes broader investigations into machine learning applications in predictive biomedicine, including genotype-based disease prediction and multimodal biomedical data integration. These works discuss challenges related to high-dimensional biomedical data, feature selection, interpretability, and integration of heterogeneous biological modalities, while highlighting the growing importance of representation learning and multimodal modeling within modern personalised medicine. As a conclusion, this dissertation demonstrates how biologically informed machine learning and representation learning approaches can support scalable predictive modeling and transcriptomics-driven therapeutic inference across diverse biomedical applications. The presented work highlights the potential of latent-space and multimodal learning frameworks to improve disease modeling, drug prioritization, and translational therapeutic prediction within personalised medicine.