Skip to content

Investigating cross-organism prediction of prokaryotic essential proteins using unsupervised language model and ensemble strategy.

Aug 2026 · Journal of Microbiological Methods · pp. 107653 · 0 citations
Medicine

TL;DR

This study proposes DeepPEP, a large language model-based framework designed to reliably transfer essential protein annotations between distantly related organisms and suggests that DeepPEP is a powerful strategy for prokaryotic essential protein prediction.

Abstract

Cross-organism prediction of essential proteins is a critical task for drug discovery and microbial engineering, yet the generalizability of existing machine learning models across diverse species remains a significant challenge. In this study, we propose DeepPEP, a large language model-based framework designed to reliably transfer essential protein annotations between distantly related organisms. Utilizing 66 curated prokaryotic datasets, we systematically evaluated DeepPEP's cross-organism performance under various conditions. Initial pairwise predictions revealed a correlation between performance and evolutionary distance; however, further investigation demonstrated that integrating training data from multiple organisms yields superior predictive power. In a benchmark scenario designed to simulate real-world applications, DeepPEP outperformed the state-of-the-art tool Geptop 2.0, showcasing a robust ability to identify species-specific essential proteins. Finally, a case study on novel genomes confirmed the model's practical effectiveness. Our results suggest that DeepPEP is a powerful strategy for prokaryotic essential protein prediction, and the rigorous evaluation framework established in this study provides a new benchmark for the field.

View source

Similar papers

Open access Aug 2026

A novel benchmark dataset for enzyme function prediction reveals the limitations of state-of-the-art models

It is demonstrated that modern EC predictors largely fail to distinguish catalytically incompetent variants from functional enzymes, and it is proposed that integrating structure-aware negative examples into both training and benchmarking is critical for developing functionally robust models in computational enzymology.

João Sartori, Ana Carolina Ramos Guimarães, Lucas de Almeida Machado · 0 citations
Open access Jul 2026

An enzyme-specific protein language model for catalytic property prediction

This manuscript introduces EnzGFM, an enzyme-specific hybrid model that improves both accuracy and efficiency across multiple prediction tasks and, together with the EnzGFM-Agent pipeline, demonstrates the ability to identify experimentally validated beneficial variants while reducing screening effort.

Chong Wang, Mengyao Li, Shaolei Geng et al. · 0 citations
Open access Jul 2026

Evaluating the cross-species transferability and scaling of sequence-to-function predictions in AlphaGenome

This work developed a scoring-approach for AI-agents to autonomously assess AlphaGenome prediction confidence and accurately differentiate between AlphaGenome’s robust sequence-level recognition across species and its current limitations when interpreting un-fine-mapped regulatory variants.

Priya Ramarao-Milne, Suyu Ma, L. Sng et al. · 0 citations
Open access Jul 2026

WASP: a pipeline for functional annotation prediction based on AlphaFold structural models

WASP highlights how structural homology can systematically discover annotations missed by sequence-based approaches, predicting protein functions from AlphaFold structures using network-based structural homology and filling metabolic model gaps by mapping 75-100% of orphan reactions.

Giorgia Del Missier, Kiyan Shabestary, Rodrigo Ledesma-Amaro · 0 citations
Conference Jul 2026

Evo-AA: Evolution-Aware Adaptation of Protein Language Models

Cutting-edge bioinformatics research is increasingly intertwined with pre-trained model techniques. However, achieving superior performance of these models in downstream applications typically requires large amounts of accurately labeled experimental data for fine-tuning, which poses substantial practical challenges due to the difficulty in preparing such datasets at scale. To address this limitation, we propose a novel few-shot fine-tuning framework, the Evolution-Aware Adaptation (Evo-AA). It aligns fine-tuning with pre-training objectives while integrating prompt learning and biological coevolutionary insights. Additionally, we introduce reinforced prompting and lambda ranking loss to further improve performance. Extensive experiments demonstrate that Evo-AA with limited training set, enhances the spearman correlation in fine-tuning tasks, while achieving superior precision and recall rates in homology search tasks. Our findings suggest that Evo-AA holds great potential to drive advancements in protein engineering and computational biology.

Yuxuan Wu, Huiqun Yu, Guisheng Fan et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.