Skip to content
Open access

A novel benchmark dataset for enzyme function prediction reveals the limitations of state-of-the-art models

Aug 2026 · bioRxiv · 0 citations · 35 references
Biology

TL;DR

It is demonstrated that modern EC predictors largely fail to distinguish catalytically incompetent variants from functional enzymes, and it is proposed that integrating structure-aware negative examples into both training and benchmarking is critical for developing functionally robust models in computational enzymology.

Abstract

Accurate computational prediction of enzyme function, standardized by Enzyme Commission (EC) numbers, is essential for large-scale genome annotation and generative enzyme design. However, it remains unclear whether state-of-the-art predictors learn the intrinsic structural determinants of catalytic activity or merely rely on global sequence similarity to annotated homologues. To address this gap, we introduce EnzymARC, a novel benchmark dataset of putative non-functional decoy sequences generated via structure-guided, systematic disruption of active sites (targeting catalytic residues and surrounding 5 Å, 10 Å, and 15 Å radii) from experimentally annotated enzymes. We evaluated three distinct prediction paradigms against this dataset: homology-based annotation (DIAMOND), contrastive learning with protein language models (CLEAN), and a deep learning model incorporating non-enzyme discrimination (DeepEC). Our findings reveal that current models are highly vulnerable to phylogenetic shortcuts. Both DIAMOND and CLEAN exhibited false positive rates exceeding 90% for low-perturbation decoys, confidently assigning the original EC numbers despite the destruction of the catalytic machinery. While DeepEC demonstrated improved sensitivity at higher perturbation levels—highlighting the benefit of negative training examples—all models struggled to identify targeted active-site disruptions. We demonstrate that modern EC predictors largely fail to distinguish catalytically incompetent variants from functional enzymes, and we propose that integrating structure-aware negative examples into both training and benchmarking is critical for developing functionally robust models in computational enzymology.

Read PDF

Similar papers

Open access Aug 2026

A Two-Stage ESM-Based Machine Learning Pipeline for Robust Hierarchical Enzyme Function Prediction

Results support the use of pretrained protein language model embeddings as an effective foundation for enzyme annotation by combining large-scale sequence representations with a lightweight supervised classifier and may facilitate functional annotation of protein sequences derived from large genomic and metagenomic datasets.

Xiao Hua, G. Grimaud · 0 citations
Aug 2026

Investigating cross-organism prediction of prokaryotic essential proteins using unsupervised language model and ensemble strategy.

This study proposes DeepPEP, a large language model-based framework designed to reliably transfer essential protein annotations between distantly related organisms and suggests that DeepPEP is a powerful strategy for prokaryotic essential protein prediction.

Ying Du, Zhikang Liu, Jing Wan · 0 citations
Open access Jul 2026

An enzyme-specific protein language model for catalytic property prediction

This manuscript introduces EnzGFM, an enzyme-specific hybrid model that improves both accuracy and efficiency across multiple prediction tasks and, together with the EnzGFM-Agent pipeline, demonstrates the ability to identify experimentally validated beneficial variants while reducing screening effort.

Chong Wang, Mengyao Li, Shaolei Geng et al. · 0 citations
Open access Jul 2026

WASP: a pipeline for functional annotation prediction based on AlphaFold structural models

WASP highlights how structural homology can systematically discover annotations missed by sequence-based approaches, predicting protein functions from AlphaFold structures using network-based structural homology and filling metabolic model gaps by mapping 75-100% of orphan reactions.

Giorgia Del Missier, Kiyan Shabestary, Rodrigo Ledesma-Amaro · 0 citations
Open access Jul 2026

Evaluating the cross-species transferability and scaling of sequence-to-function predictions in AlphaGenome

This work developed a scoring-approach for AI-agents to autonomously assess AlphaGenome prediction confidence and accurately differentiate between AlphaGenome’s robust sequence-level recognition across species and its current limitations when interpreting un-fine-mapped regulatory variants.

Priya Ramarao-Milne, Suyu Ma, L. Sng et al. · 0 citations
Open access Jul 2026

Minimal Data · Maximal Insight (MDMI): A Structure-guided Pipeline for Discovering Functional Alternatives in Peptide-Protein Interfaces

Minimal Data Maximal Insight (MDMI), a two-stage structure-guided computational pipeline that designs functional peptide variants using only a small, annotated dataset, demonstrates that structure-informed pipelines can uncover remote functional sequence space from minimal data.

P. Bayat, Spencer J. Perkins, Sebastian Clancy et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.