A Novel Unsupervised Learning Model with Bayesian Hyperparameter Optimization for Predicting Substrates of Promiscuous Enzymes
Abstract
Promiscuous enzymes catalyze multiple biochemical reactions, but predicting their substrate profiles remains challenging because annotations are incomplete, reliable negative labels are scarce, and enzyme–substrate relationships are inherently multilabel. Here, we present PreSEPM (Predictive Substrate Explorer for Promiscuous Enzymes), a sequence-only, closed-set framework that combines pretrained protein representations, Gaussian mixture modeling, cluster–substrate alignment, and Bayesian optimization of decision thresholds. PreSEPM operates within a predefined substrate panel and does not require explicit substrate or structural descriptors. On a UniProt-derived triacylglycerol lipase data set, PreSEPM achieved an AUROC of 0.86, an AUPRC of 0.73, and a maximum F1 score of 0.71, outperforming the evaluated conventional machine learning baselines. On an independent high-throughput lipase screen, PreSEPM also outperformed a strong single-task logistic-regression baseline under enzyme-wise cross-validation, increasing Macro-AUPRC from 0.65 to 0.68 and F1 max from 0.63 to 0.69. Under the cross-data set setting, SMOTE-based augmentation increased AUPRC from 0.62 to 0.70. These results support PreSEPM as a useful sequence-based baseline for substrate profile completion and experimental prioritization within promiscuous enzyme families when annotations are sparse and structural or ligand-level information is unavailable.