Skip to content
Open access

A Two-Stage ESM-Based Machine Learning Pipeline for Robust Hierarchical Enzyme Function Prediction

Aug 2026 · bioRxiv · 0 citations · 60 references
Biology

TL;DR

Results support the use of pretrained protein language model embeddings as an effective foundation for enzyme annotation by combining large-scale sequence representations with a lightweight supervised classifier and may facilitate functional annotation of protein sequences derived from large genomic and metagenomic datasets.

Abstract

Accurate enzyme annotation remains a major bottleneck in translating rapidly growing protein sequence data into biological knowledge. Enzyme Commission (EC) prediction is particularly challenging because enzyme functions are organized hierarchically, annotations are often imbalanced across classes, and sequence similarity alone may be insufficient to resolve functional differences. To address these challenges, we developed ESM-ECForest, a two-stage framework that combines protein embeddings generated by the pretrained language model ESM-2 (Evolutionary Scale Modeling 2) with Random Forest classifiers. The first stage distinguishes enzymes from non-enzymes, whereas the second assigns one or more EC numbers to proteins predicted to be enzymatic. On an external benchmark comprising 25,778 protein sequences, ESM-ECForest achieved the highest weighted F1 score among the evaluated methods at all four EC levels, decreasing from 0.94 at Level 1 to 0.90 at Level 4. The largest relative improvements were observed for lyases (EC 4), ligases (EC 6), and translocases (EC 7), although EC 6 and EC 7 remained the most difficult classes internally. Visualization of the ESM-2 embedding space using Uniform Manifold Approximation and Projection (UMAP) revealed clustering patterns consistent with enzyme functional relationships, indicating that biologically relevant information is retained in the pretrained representations prior to supervised classification. These results support the use of pretrained protein language model embeddings as an effective foundation for enzyme annotation. By combining large-scale sequence representations with a lightweight supervised classifier, ESM-ECForest provides a scalable approach for EC prediction and may facilitate functional annotation of protein sequences derived from large genomic and metagenomic datasets.

Read PDF

Similar papers

Open access Jul 2026

An enzyme-specific protein language model for catalytic property prediction

This manuscript introduces EnzGFM, an enzyme-specific hybrid model that improves both accuracy and efficiency across multiple prediction tasks and, together with the EnzGFM-Agent pipeline, demonstrates the ability to identify experimentally validated beneficial variants while reducing screening effort.

Chong Wang, Mengyao Li, Shaolei Geng et al. · 0 citations
Open access Aug 2026

FuncSeek: Multi-PLM contrastive learning for protein functional similarity search

FuncSeek is described, a contrastive learning model which utilizes three diverse, complementary PLMs: ESM2 (to model evolutionary co-variation), ProstT5 (for bilingual sequence and structure embeddings), and ProteinBERT (for functional semantic similarities) that each capture a different aspect of protein biology: evolutionary patterns, three-dimensional shape, and functional context.

Leendert J. Cloete, Hugh G. Patterton · 0 citations
Open access Aug 2026

A novel benchmark dataset for enzyme function prediction reveals the limitations of state-of-the-art models

It is demonstrated that modern EC predictors largely fail to distinguish catalytically incompetent variants from functional enzymes, and it is proposed that integrating structure-aware negative examples into both training and benchmarking is critical for developing functionally robust models in computational enzymology.

João Sartori, Ana Carolina Ramos Guimarães, Lucas de Almeida Machado · 0 citations
Open access Aug 2026

Data-Centric Evaluation of Protein Function Prediction Pipelines

Findings show that performance estimates in protein function prediction should be interpreted as outcomes of complete data-centric workflows rather than isolated properties of predictive models.

Nicole Soto-García, Norma Murillo-Acevedo, Julián García-Vinuesa et al. · 0 citations
Open access Sep 2026

Squidly harnesses enzyme functional hierarchy and contrastive learning to efficiently predict catalytic residues from sequence

Enzymes present a sustainable alternative to traditional chemical industries, drug synthesis, and bioremediation applications. Because catalytic residues are the key amino acids that drive enzyme function, their accurate prediction facilitates enzyme function prediction. Sequence similarity-based approaches such as BLAST are fast but require previously annotated homologues. Machine-learning (ML) approaches aim to overcome this limitation; however, current gold-standard ML-based methods require high-quality 3D structures limiting their application to large datasets. To address these challenges, we developed Squidly, a sequence-only tool that leverages contrastive representation learning with a biology-informed, rationally designed pairing scheme to distinguish catalytic from non-catalytic residues using per-token Protein Language Model embeddings. Squidly surpasses state-of-the-art ML annotation methods in catalytic residue prediction while remaining sufficiently fast to enable wide-scale screening of databases. We ensemble Squidly with BLAST to provide an efficient tool that annotates catalytic residues with high precision and recall for both in- and out-of-distribution sequences.

W. J. Rieger, Mikael Bodén, Frances H. Arnold et al. · 0 citations
Open access Jul 2026

MAERM: Predicting Enzyme-Reaction Matching Relationships with a Mixed-Attention Model

Harnessing enzyme specificity requires a thorough understanding of enzyme promiscuity, which determines enzymes’ catalytic scope; however, measuring this scope still relies heavily on labor-intensive analytical approaches. While data-driven approaches have emerged to predict the catalytic scope of enzymes, these methods continue to face challenges such as restricted datasets and insufficient integration of enzyme structural information and reaction transformations. Here, we introduce MAERM, an innovative mixed-attention model designed to predict enzyme-reaction matching relationships. Built on our MAERM-DB, a dataset with broad coverage of validated and chemoenzymatic catalysis data, MAERM utilizes a local-global attention module to integrate multimodal enzyme information with fine-grained reaction representations, thereby predicting enzyme-reaction matching probabilities. Results show that MAERM consistently outperforms all baselines, with an average F1-score of 0.984. Notably, on challenging test samples with less than 40% sequence identity to the training set, MAERM outperforms the second-ranked model by 5.9% in F1-score. In addition, MAERM achieves the highest top-10 success rate of 51.7% on Enzyme-405 and the highest balanced accuracy of 0.697 on BioCat-547, further supporting its generalizability in enzyme screening and chemoenzymatic catalysis. Finally, MAERM can serve as an efficient scoring module. When integrated with ProteinMPNN, MAERM has successfully guided novel enzyme design for two carbonyl reduction reactions, resulting in enhanced catalytic potential for the native substrate and demonstrating broad compatibility. Overall, MAERM has the potential to reduce the experimental cost of measuring enzymes’ catalytic scope, facilitate enzyme design, and ultimately accelerate the design-build-test-learn cycle in enzyme engineering.

Tiantao Liu, Silong Zhai, Shaolong Lin et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.