Aug 2026· International Journal of Machine Learning and Cybernetics· Vol 17· 0 citations· 36 references
TL;DR
A novel model SSIPIN is proposed for protein function prediction, using three types of data including protein sequences, protein-protein interaction networks (PPIN) and protein tertiary structures, which outperforms existing methods on multiple datasets.
SPPIPred, an advanced machine learning-based model designed for precise PPI prediction, is presented, offering valuable insights to researchers in the field of bioinformatics and improving applications within bioengineering and pharmaceutical development.
M. Rahman, M. Ali, Md. Shohidullah et al.· PLoS ONE· 0 citations
Protein–protein interactions (PPIs) play a crucial role in enabling proteins to carry out their functions within various biological processes (Hui et al., 2003). Since the introduction of the yeast two-hybrid (Y2H) method for PPI detection in 1989 (Fields and Song, 1989), the identification of PPIs has become a significant focus in modern biological research. PPI goes beyond examining individual proteins, allowing researchers to establish a comprehensive network that regulates biological processes. Rice, as a key model organism in plant biological studies, has been at the forefront of PPI research. In 2008, prominent rice scientists in China called for concerted efforts to define a comprehensive protein–protein interaction network experimentally, which aimed to facilitate the prediction of the functional mechanisms operating throughout a plant’s lifecycle (Zhang et al., 2008). With efforts for 2 decades, the experimentally identified rice PPIs have reached over ten thousand. Several public databases have been established to systematically collate and store PPIs, including STRING (Szklarczyk et al., 2019), BioGRID (Oughtred et al., 2020), IntAct (del Toro et al., 2022), PRIN (Gu et al., 2011), RicePPINet (Liu et al., 2017) and RiceNet v2 (Lee et al., 2015). However, most PPI datasets in rice stem from computational predictions, while experiment-based rice PPI datasets are fragmented due to the lack of systematic profiling at the rice PPIome level, which largely hinders information sharing in the rice research community. To bridge this gap, we constructed the Port of Protein-Protein Interactomes (POPPIN; https://riceome.hzau.edu.cn/poppin/), an integrated database dedicated to sharing experimentally verified PPIs and functional clues in rice. Empowered by high-throughput PPIome profiling technologies and text mining assisted by a large language model (Huang et al., 2025; Liu et al., 2025), POPPIN currently has deposited over 150,451 pieces of rice PPI-related information. Additionally, POPPIN provides detailed protein information, including GO annotations, subcellular localizations, domains, trait ontology (TO) information, and hyperlinks to external biological databases. Through offering a user-friendly web interface for search and dynamic network visualization, POPPIN serves as the first large-scale, experiment-based database for searchable PPIs in rice, and has the potential to be extended to other species under this structural framework.
The fundamental relationship among protein sequence, structure, function, and physicochemical properties is a central principle in biology. While in principle protein function and properties should be able to be derived directly from protein sequence, in practice protein function and property prediction methods have been designed around specific datasets and specific property or function subsets, leading to an enormous gap between function annotation and property prediction. To address these challenges, we introduce AlphaFunctor, a category theory based foundation model-like platform to bridge the gap between protein function annotation and property prediction. Based on the hypothesis that protein function and properties can be directly derived from protein sequence, AlphaFunctor predicts protein functions as represented by Gene Ontology terms directly from sequence. Using these function predictions, AlphaFunctor further maps protein functions using topological spectral theory, path-complex neural networks, and protein domain analysis onto downstream property prediction. AlphaFunctor is (pre)trained in nearly 0.6 million protein function data points to deliver the state-of-the-art protein function annotation on three benchmark datasets. Without task-specific network redesign, AlphaFunctor maps qualitative protein function annotation to various qualitative and quantitative protein property predictions, outperforming other dataset-specific and task-specific competing predictors.
Xiang Liu, Anna E. Yee, J. Vermaas et al.· arXiv.org· 0 citations
A neural network-based pipeline that integrates amino acid sequences with structural features is developed and provides a modular prototype for follow-up, more extensive protein modeling, including larger proteins and sequence of variable sizes.
Carl David Jasper Causin, M. Fyta· APL Machine Learning· 0 citations
Whether AlphaFold 3 complex prediction, combined with STRING evidence and domain-level analysis of interfaces and interaction partners, can help identify and characterize DUF-containing proteins and suggest roles for DUF4130 in nucleic-acid-associated radical-SAM biology and DUF5819 in a bacterial system related to vitamin-K-dependent carboxylation are suggested.
Lino Riepenhausen, Francesco Costa, Antonina Andreeva et al.· bioRxiv· 0 citations
The autoencoder framework encodes protein sequence information related to domains, families, and patterns—into a lengthy, sparse binary vector that outperforms other neural network models, including convolutional neural networks, recurrent neural networks, long short-term memory networks, and bidirectional long short-term memory networks.
Biswajit Senapati, Ranjita Das· Journal of Computer-Aided Mo...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.