Enzymes present a sustainable alternative to traditional chemical industries, drug synthesis, and bioremediation applications. Because catalytic residues are the key amino acids that drive enzyme function, their accurate prediction facilitates enzyme function prediction. Sequence similarity-based approaches such as BLAST are fast but require previously annotated homologues. Machine-learning (ML) approaches aim to overcome this limitation; however, current gold-standard ML-based methods require high-quality 3D structures limiting their application to large datasets. To address these challenges, we developed Squidly, a sequence-only tool that leverages contrastive representation learning with a biology-informed, rationally designed pairing scheme to distinguish catalytic from non-catalytic residues using per-token Protein Language Model embeddings. Squidly surpasses state-of-the-art ML annotation methods in catalytic residue prediction while remaining sufficiently fast to enable wide-scale screening of databases. We ensemble Squidly with BLAST to provide an efficient tool that annotates catalytic residues with high precision and recall for both in- and out-of-distribution sequences.
W. J. Rieger, Mikael Bodén, Frances H. Arnold et al.· eLife· 0 citations
It is shown that integrating genome-informed data augmentation with contrastive learning in protein language space enables identification of phenazine-interacting proteins starting from only 14 known phenazine modifying sequences, and coupling comparative genomics with protein machine learning can convert “small data” typically outside the scope of machine learning into actionable predictive power, thereby facilitating enzyme discovery.
Xiaoyu Shan, I. Trindade, N. Glasser et al.· Proceedings of the National...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.