ERAM aligns pre-trained molecular representations from Protein Language Model with the knowledge of enzyme catalysis by modeling enzymatic reactions as multi-relational data, and demonstrates its potential as a versatile and effective tool for enzyme catalysis research.
Abstract
Enzymatic reactions play an emerging role in a broad spectrum of scientific and industrial applications. The inherent complexity of enzymes, such as their substrate specificity, conformational flexibility, and the vast diversity of reactions involved, poses substantial challenges for the advanced computational prediction of enzymatic reactions with desirable accuracy. Moreover, existing approaches are mostly tailored for a specific sub-task, such as substrate prediction or binding site annotation, which limits their applicability. In this study, we introduce ERAM, a task-agnostic multimodal learning framework capable of addressing a broad range of downstream applications with both accuracy and efficiency. ERAM aligns pre-trained molecular representations from Protein Language Model with the knowledge of enzyme catalysis by modeling enzymatic reactions as multi-relational data. In enzyme retrieval tasks, ERAM achieves an improvement of 28.31% in mean average precision compared with the state-of-the-art (SOTA) method, CREEP. In substrate prediction tasks, ERAM outperforms the SOTA method ESP, achieving average improvements of 35.53% and 22.97% in Matthews correlation coefficient across two datasets. Additionally, ERAM exhibits commendable interpretability by assigning higher attention weights to binding sites, resulting in lower false-positive rates (42.36%) and higher overlap scores (70.59%) in the unsupervised binding site prediction task compared to RXNAA Mapper. By learning embeddings of substrates, enzymes, and products within a unified knowledge graph latent space, ERAM demonstrates its potential as a versatile and effective tool for enzyme catalysis research.
Acidophilic proteins that remain stable and functional under highly acidic conditions, are important for industrial biocatalysis, acid-related bioprocessing, and the discovery of acid-stable enzymes. However, their identification relies heavily on time-consuming experimental screening methods. With the rapid growth of protein sequence databases, the need for computational identification methods that are both accurate and efficient has become stronger. The emergence of protein language models (PLMs) has significantly improved the sequence representation of downstream biological prediction tasks. This paper proposes MI-PEFT, a mixture-of-experts integrated parameter-efficient fine-tuning framework. Built on the ESM C-600M backbone, the framework incorporates LoRA-based PEFT methods and a DeepSeekMoE-based classification head to resolve the limitations of PEFT and significantly improve computational efficiency. Notably, this task is characterized by a significant class imbalance in the dataset, making high specificity particularly challenging. The experimental results demonstrate that MI-PEFT on PLMs, especially {\text{C}}^{\text{3}}\text{A}, serves as an efficient tool for identifying acidophilic proteins and a constrained pathway that helps resolve class-imbalance by preserving the pretrained representations.
This manuscript introduces EnzGFM, an enzyme-specific hybrid model that improves both accuracy and efficiency across multiple prediction tasks and, together with the EnzGFM-Agent pipeline, demonstrates the ability to identify experimentally validated beneficial variants while reducing screening effort.
The enzyme kinetic parameters, including the turnover number, Michaelis constant, and inhibition constant, are key metrics for assessing catalytic performance. Although deep learning models have recently incorporated multimodal information from enzymes and substrates to predict these parameters, several obstacles still persist. First, current data sets suffer from limited size, inconsistency, and a lack of unified standards. Second, most existing approaches prioritize cross-modal consistency but fail to sufficiently exploit the unique information residing in each individual modality. Meanwhile, although a limited number of studies have recognized that collaborative exploration of shared and specific information can enhance model performance, these methods remain difficult to directly apply to enzyme–substrate pairs, as enzyme–substrate relationships are inherently interactive rather than semantically equivalent counterparts. Third, the measured kinetic parameters are often unevenly distributed, which severely undermines the predictive accuracy of existing models when dealing with extreme value ranges. To resolve the above challenges, we first compile Kinetic-DB, a large-scale and consistently formatted data set from public resources. Building upon this data set, we develop GAPEK, a new framework for estimating enzyme kinetic parameters. In particular, an adaptive data augmentation module is devised to enrich the diversity of both enzyme and substrate sequences, thereby alleviating the adverse effects of data imbalance. Subsequently, we perform feature extraction using two pretrained models, ESM-2 for enzymes and Mole-BERT for substrates, to obtain multimodal embeddings. To decouple these complex interacting features, we introduce a tailored dual information exploration module to capture both modality-specific and cross-modal information, further refined by domain classification and distribution alignment loss functions. To explicitly handle the imbalanced data distribution, our base model, GAPEK, incorporates an adaptive density-weighted loss function. Building on this, we propose GAPEK+, which integrates the Squared Error Relevance Area (SERA) function to reconfigure the learning objective. By prioritizing high-relevance regions, GAPEK+ effectively calibrates the model’s sensitivity to rare but critical extreme values, substantially mitigating the prediction bias inherent in heavy-tailed regression tasks. Experimental results demonstrate that both GAPEK and GAPEK+ achieve superior performance over existing state-of-the-art approaches, particularly across extreme parameter ranges, highlighting their potential as valuable tools applicable to enzyme engineering, synthetic biology, and drug discovery.
Cheng-Hao Zhu, Wei-Ping Ding, Wei Zhang et al.· Journal of Chemical Informat...· 0 citations
Identifying enzymes capable of catalyzing specific chemical transformations across large sequence databases remains a major challenge in biocatalyst discovery. Conventional fingerprint-based methods capture global molecular structure but fail to represent bond-breaking and bond-forming events, limiting generalization to structurally novel reactions. We introduce a dual-track evaluation framework to distinguish true generalization from memorization, assessing retrieval on structurally isolated reactions (n = 50) within a 63,259-sequence enzyme pool. The strongest fingerprint baseline achieves R@10 = 0.020. To address this limitation, we develop GATv2-ECR, a heterogeneous dual-tower model integrating reaction-center graph encoding, a frozen ESM-2 sequence encoder, contrastive learning, and EC-aware soft reranking. GATv2-ECR achieves R@10 = 0.160 on isolated queries and R@10 = 0.308 on an out-of-distribution subset (n = 39), capturing mechanistically relevant features and supporting generalizable enzyme retrieval under open-world conditions.
Unknown authors· Journal of Chemical Informat...· 0 citations
Enzymes present a sustainable alternative to traditional chemical industries, drug synthesis, and bioremediation applications. Because catalytic residues are the key amino acids that drive enzyme function, their accurate prediction facilitates enzyme function prediction. Sequence similarity-based approaches such as BLAST are fast but require previously annotated homologues. Machine-learning (ML) approaches aim to overcome this limitation; however, current gold-standard ML-based methods require high-quality 3D structures limiting their application to large datasets. To address these challenges, we developed Squidly, a sequence-only tool that leverages contrastive representation learning with a biology-informed, rationally designed pairing scheme to distinguish catalytic from non-catalytic residues using per-token Protein Language Model embeddings. Squidly surpasses state-of-the-art ML annotation methods in catalytic residue prediction while remaining sufficiently fast to enable wide-scale screening of databases. We ensemble Squidly with BLAST to provide an efficient tool that annotates catalytic residues with high precision and recall for both in- and out-of-distribution sequences.
W. J. Rieger, Mikael Bodén, Frances H. Arnold et al.· eLife· 0 citations
Promiscuous enzymes catalyze multiple biochemical reactions, but predicting their substrate profiles remains challenging because annotations are incomplete, reliable negative labels are scarce, and enzyme–substrate relationships are inherently multilabel. Here, we present PreSEPM (Predictive Substrate Explorer for Promiscuous Enzymes), a sequence-only, closed-set framework that combines pretrained protein representations, Gaussian mixture modeling, cluster–substrate alignment, and Bayesian optimization of decision thresholds. PreSEPM operates within a predefined substrate panel and does not require explicit substrate or structural descriptors. On a UniProt-derived triacylglycerol lipase data set, PreSEPM achieved an AUROC of 0.86, an AUPRC of 0.73, and a maximum F1 score of 0.71, outperforming the evaluated conventional machine learning baselines. On an independent high-throughput lipase screen, PreSEPM also outperformed a strong single-task logistic-regression baseline under enzyme-wise cross-validation, increasing Macro-AUPRC from 0.65 to 0.68 and F1 max from 0.63 to 0.69. Under the cross-data set setting, SMOTE-based augmentation increased AUPRC from 0.62 to 0.70. These results support PreSEPM as a useful sequence-based baseline for substrate profile completion and experimental prioritization within promiscuous enzyme families when annotations are sparse and structural or ligand-level information is unavailable.
Rong-Sheng Gao, Cheng-Ye Duan, Nan Qin et al.· ACS Omega· 0 citations
What does it take to trust AI-driven HVAC optimization? Our AI Model Factory combines agents, machine learning, reinforcement learning and deterministic checks in a governed workflow designed for messy, real-world building data. The post We built an AI factory for HVAC control appeared first on GPT-Lab.
MIT News · Artificial Intelligence· news.mit.eduAug 27, 2026
A new machine-learning framework aims to improve the success rate of computational protein design while moving away from results that reproduce sequences found in nature.
MIT News · Artificial Intelligence· news.mit.eduAug 18, 2026
A new method for surgically removing training examples from a model reveals that as datasets grow, the link between what a model learns and what it produces dissolves.
Microsoft Research Blog· microsoft.comJul 30, 2026
LLMs do not get smarter just by remembering more. EvoLib turns experience into evolving knowledge, taking reusable skills and insights that help models learn and adapt across tasks long after deployment. The post EvoLib: Turning experience into evolving knowledge appeared first on Microsoft Research.
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.