Skip to content
Open access

FuncSeek: Multi-PLM contrastive learning for protein functional similarity search

Aug 2026 · bioRxiv · 0 citations · 1 references
Biology

TL;DR

FuncSeek is described, a contrastive learning model which utilizes three diverse, complementary PLMs: ESM2 (to model evolutionary co-variation), ProstT5 (for bilingual sequence and structure embeddings), and ProteinBERT (for functional semantic similarities) that each capture a different aspect of protein biology: evolutionary patterns, three-dimensional shape, and functional context.

Abstract

Below 30% pairwise sequence identity, alignment-based methods struggle to reliably distinguish true homologs from chance, and enzyme function prediction degrades accordingly: on proteins in this regime, even advanced methods (CLEAN) achieves only 55.1% accuracy at full EC specificity on the CARE benchmark. To this end Protein Language Models (PLMs) have gained favor as alternatives. However, PLMs often encode only a subset of the biology, whereas the understanding of enzyme function requires among other things a combination of sequence, structure and functional-context simultaneously. In this work, we describe FuncSeek, a contrastive learning model which utilizes three diverse, complementary PLMs: ESM2 (to model evolutionary co-variation), ProstT5 (for bilingual sequence and structure embeddings), and ProteinBERT (for functional semantic similarities). Using SwissProt data, these 2816-D embeddings are labelled with Enzyme Commission numbers (EC) and are trained through a supervised contrastive head into a 256-D space. FuncSeek attains 64.6% nearest-neighbour EC4 accuracy on the CARE out-of-distribution benchmark set (ood30; proteins below 30% identity to training set), outperforming CLEAN (55.1%) and Diamond BLASTp (51.4%), and obtains 93.7% nearest-neighbour EC4 accuracy on the promiscuous, multi-functional enzymes benchmark (CLEAN, 69.4%). We also show that the learned representations transfer without retraining to the TrEMBL database, achieving 97.3% nearest-neighbour EC4 accuracy on an 8,031 BRENDA-validated enzyme set, never seen during training. Because only projected embeddings are stored in the target index, and function is inferred from an annotated reference set, we propose this paradigm for rapidly searching extremely large metagenomic databases, bypassing costly sequence alignment and annotation pipelines. Author Summary Enzymes are the molecular machines that carry out the chemistry of life, and knowing which reaction an enzyme performs is essential for medicine, biotechnology, and understanding how organisms work. Yet new protein sequences are being discovered far faster than we can study them in the laboratory, and the most common computational shortcut, comparing a new sequence to well-studied ones, breaks down when the sequences are only distantly related. We asked whether recent artificial intelligence models that learn the “language” of proteins could close this gap. Rather than rely on a single model, we combined three that each capture a different aspect of protein biology: evolutionary patterns, three-dimensional shape, and functional context. We then trained a system, which we call FuncSeek, to arrange enzymes so that those performing the same reaction sit close together. FuncSeek predicted enzyme function more accurately than the leading existing tools, especially for distantly related and multi-functional enzymes, and this accuracy carried over to proteins it had never seen. Because it represents each protein as a compact numerical fingerprint, it can search enormous, unexplored collections of sequences quickly.

Read PDF

Similar papers

Open access Aug 2026

PLMView: collaborative protein language model representations for fast and scalable specialized protein function inference

Applications to thioredoxins, visual opsins, and Tara Oceans environmental diatom cold-shock proteins show that PLMView can move from interpretable residue-level determinants in well-studied protein families to large-scale environmental functional discovery, linking molecular specialization to ecological distribution and transcriptional deployment across the global ocean.

Vinh-Son Pho, Alessandro Natale Bianchi, Mattéo Scarsini et al. · 0 citations
Open access Aug 2026

A Two-Stage ESM-Based Machine Learning Pipeline for Robust Hierarchical Enzyme Function Prediction

Results support the use of pretrained protein language model embeddings as an effective foundation for enzyme annotation by combining large-scale sequence representations with a lightweight supervised classifier and may facilitate functional annotation of protein sequences derived from large genomic and metagenomic datasets.

Xiao Hua, G. Grimaud · 0 citations
Open access Jul 2026

TEDlm: domain-centric protein language models with optional structural pre-training

TEDlm variants also substantially improve zero-shot Molecular Function prediction over ESM2, while matching it on various biophysical property tasks, indicating that signals are largely domain-intrinsic.

Tiejun Wei, S. Kandathil, Daniel W. A. Buchan et al. · 0 citations
Open access Aug 2026

Data-Centric Evaluation of Protein Function Prediction Pipelines

Findings show that performance estimates in protein function prediction should be interpreted as outcomes of complete data-centric workflows rather than isolated properties of predictive models.

Nicole Soto-García, Norma Murillo-Acevedo, Julián García-Vinuesa et al. · 0 citations
Open access Jul 2026

Predictions of protein–protein interactions: Learning sequences and structures

A neural network-based pipeline that integrates amino acid sequences with structural features is developed and provides a modular prototype for follow-up, more extensive protein modeling, including larger proteins and sequence of variable sizes.

Carl David Jasper Causin, M. Fyta · 0 citations
Open access Aug 2026

ProtJEPA: A Multimodal Joint-Embedding Predictive Architecture for Protein Biological World Modeling with Multi-Teacher Modality-Attentive Fusion

Over 99.9% of known protein sequences lack experimentally validated functional annotations. We present ProtJEPA, a multimodal Joint-Embedding Predictive Architecture that trains a sequence-only student encoder to predict joint embeddings spanning ten biological modalities—sequence, structure, knowledge graph, protein interactions, literature, localization, tissue expression, GO function, anatomy, and disorder—requiring only sequence at inference. The key innovation is target whitening, which eliminates severe anisotropy in joint targets (mean cosine 0.984 to 0.086) and prevents representation collapse without covariance regularization. On 1,828 held-out dark proteins with zero primary Pfam family overlap with training, ProtJEPA achieves 58.07% Hit@10 on zero-shot GO retrieval (+2.80 pp, p = 0.020), 69.99% enzyme class accuracy (+9.64 pp, p < 0.001), and +11.87 pp subcellular localization at 1% labels (p < 0.001). Under realistic dark-protein deployment conditions where relational modalities are unavailable, ProtJEPA significantly outperforms naive concatenation of remaining modalities. Cross-domain evaluations on drug–target interaction and disorder prediction confirm transfer beyond training modalities, with the T1-only < ESMC < ProtJEPA ordering replicated across six independent tasks. Ablations establish that Phase 1 aggregator pretraining and target whitening are each independently load-bearing.

Vaibhava Lakshmi Ravideshik, Jinha Kim, M. Kellis · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.