Skip to content
Open access

S2DA-GO: enhancing protein function prediction via gradient-decoupled cross-attention and semantic priors

Jul 2026 · Frontiers in Genetics · Vol 17 · 0 citations · 36 references
Medicine

TL;DR

S2DA-GO alleviates feature interference and improves prediction for sparsely annotated GO terms, providing a promising framework for large-scale annotation of uncharacterized proteins.

Abstract

Accurate automated protein function prediction is essential for bridging the widening gap between the exponential accumulation of uncharacterized protein sequences and the limited repository of experimentally verified functional annotations. However, despite recent advances in incorporating Gene Ontology (GO) priors and sequence-label interactions, existing computational methods still face substantial challenges in long-tailed multi-label settings, particularly in handling optimization instability caused by rare-label noise and in learning reliable representations for sparsely annotated GO terms. To address these challenges, we propose S2DA-GO, a sequence-based model using pre-trained protein language model embeddings and GO textual semantic priors for protein function prediction. S2DA-GO integrates a global contextual stream with a local target-aware stream to capture multi-scale functional patterns. To alleviate optimization instability caused by long-tailed label noise, we introduce a Gradient-Decoupled Cross-Attention Module (GDCAM), which reduces the interference of label-specific gradients on the shared backbone. In addition, we incorporate learnable residual semantic priors derived from BioBERT-encoded GO definitions, enhancing the model’s adaptability to rare functional terms. On the benchmark dataset, S2DA-GO outperformed the strong baseline GDTGO across all three GO branches, achieving notable relative AUPR improvements of 4.0% Molecular Function (MF), 6.0% Biological Process (BP), and 8.6% Cellular Component (CC), with AUPR scores reaching 65.1%, 33.4%, and 41.7%, respectively. Notably, S2DA-GO remained robust in low-homology settings and provided interpretable residue-level signals that accurately aligned with experimentally verified binding sites. Overall, S2DA-GO alleviates feature interference and improves prediction for sparsely annotated GO terms, providing a promising framework for large-scale annotation of uncharacterized proteins.

Read PDF

Similar papers

Preprint Aug 2026

Interpreting Latent Protein Language Model Features with Geometric Annotations

This work introduces an automated and scalable method for interpreting SAE features in ESM-2 by using geometrically inspired features of the protein $\text{C}_\alpha$ backbone, providing a robust method of annotating proteins activated within SAE neurons at a residue level, providing a bridge between mechanistic interpretability and structural biology.

S. Setlur, Djordje Mihajlovic, Darrick Lee · 0 citations
Aug 2026

HiHPO: Multimodal Hierarchical Graph Learning for Predicting Missing Protein-Phenotype Associations.

Understanding protein-phenotype associations is essential for elucidating disease mechanisms and supporting phenotype-driven diagnosis. Although the Human Phenotype Ontology (HPO) provides a standardized framework for phenotypic description, protein-HPO annotations remain incomplete and continuously evolving, posing challenges for robust computational prediction. Existing methods often fail to fully exploit hierarchical phenotype semantics and multimodal biological context, particularly under sparse annotation settings. We propose HiHPO, a multimodal, hierarchy-aware graph contrastive learning framework for predicting protein-phenotype associations. HiHPO integrates complementary biological information from protein-protein interaction networks, gene expression profiles, and protein language model embeddings, while explicitly incorporating HPO hierarchical structure into contrastive representation learning. This design enables the model to preserve semantic relationships among phenotypes and improve generalization to fine-grained and sparsely annotated terms. Extensive evaluations on both random and temporal validation splits demonstrate that HiHPO consistently outperforms state-of-the-art methods, with pronounced advantages on deep HPO terms and newly curated annotations. Additional analyses confirm the contribution of each modality and the robustness of the framework across varying annotation densities. These results highlight the potential of hierarchy-aware multimodal learning for advancing protein-phenotype association prediction and disease-related biomedical research. Code and data are available at https://github.com/ZhuLab-Fudan/HiHPO.

Hancheng Liu, W. Zhai, Shaojun Wang et al. · 0 citations
Open access Aug 2026

FuncSeek: Multi-PLM contrastive learning for protein functional similarity search

FuncSeek is described, a contrastive learning model which utilizes three diverse, complementary PLMs: ESM2 (to model evolutionary co-variation), ProstT5 (for bilingual sequence and structure embeddings), and ProteinBERT (for functional semantic similarities) that each capture a different aspect of protein biology: evolutionary patterns, three-dimensional shape, and functional context.

Leendert J. Cloete, Hugh G. Patterton · 0 citations
Open access Aug 2026

Attention-Enhanced Bi-LSTM Ensembles with Frozen ESM-2 Embeddings Achieve Competitive Performance in Protein Subcellular Localization

Attention weight analysis shows the model identifies biologically relevant signals, including C-terminal membrane regions and N-terminal mitochondrial sequences, confirming that ESM-2 embeddings enable effective performance with simplified architectures.

Johaimen Omar · 0 citations
Open access Jul 2026

X-PAIR: an ultrafast multitask framework for proteome-scale reconstruction of PPI networks and partner-specific interfaces from sequence

X-PAIR is presented, a sequence-based multitask deep learning framework that jointly predicts whether two proteins interact and identifies their partner-specific interface residues, and links proteome-scale interaction discovery to the residue-level determinants of partner-specific molecular recognition.

S. Rescalli, A. Carbone · 0 citations
Open access Aug 2026

ETAP-CLF: an ESM3-based transformer attention framework for binary protein classification

Binary protein classification supports diverse tasks in computational biology, including pathway-membership inference and sequence-based candidate prioritization. Protein language models generate information-rich residue-level representations, but downstream classifiers commonly compress them using fixed pooling operations that may discard task-relevant sequence context. We present ETAP-CLF, a compact framework that combines pretrained per-residue ESM3 embeddings with lightweight transformer contextualization and learned attention pooling to classify variable-length proteins and generate residue-level attention scores. The ESM3 parameters remained frozen, and the same ETAP-CLF architecture and hyperparameter configuration were used across ferroptosis-, senescence-, and pyroptosis-associated protein prediction. ETAP-CLF achieved AUROCs of 0.98, 0.95 and 0.91 for these tasks, respectively. In the ferroptosis benchmark, ETAP-CLF outperformed the evaluated published models. These results demonstrate that a common downstream design can adapt to multiple process-associated classification tasks without fine-tuning the task-specific model architecture. ETAP-CLF provides a generalizable approach for sequence-based protein prioritization and a basis for broader evaluation across binary protein-classification problems.

Jianyu Ren, Hanli Jiang, Puchangxin Li et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.