KRstereo: Predicting β-Hydroxy Stereochemistry in Polyketides Using Protein Language Models
Abstract
Polyketides are a major class of bioactive natural products whose activities are often determined by the stereochemistry of β-hydroxyl groups. In type I polyketide synthases (PKSs), ketoreductase (KR) domains establish these stereocenters, making accurate prediction of KR stereochemistry important for natural product discovery and PKS engineering. Existing rule-based methods rely on a limited set of sequence motifs and often perform poorly across phylogenetically diverse taxa. Here, we present KRstereo, a machine learning framework that predicts KR stereochemistry directly from sequence. Analysis of β-modular KR domains from MIBiG 3.1 identified informative sequence features beyond canonical motifs, motivating the use of protein language model embeddings. KRstereo achieved accuracies of up to 95.7% across taxa and 93.0% for non-Streptomyces KRs, consistently outperforming existing rule-based approaches. Validation using newly characterized KR domains from MIBiG 4.0 confirmed strong generalizability, including cases misclassified by current methods. Application of KRstereo to 20,840 β-modular KR domains from antiSMASH enabled large-scale stereochemical annotation of previously uncharacterized PKS systems. By linking sequence to stereochemical function, KRstereo improves reconstruction of polyketide structures from biosynthetic gene clusters and facilitates stereochemistry-aware genome mining and engineering of PKS assembly lines.