Aug 2026· IEEE journal of biomedical and health informatics· Vol PP· 0 citations
Medicine
TL;DR
Results validate PreMemMoRF as a robust and reliable computational framework for the large-scale identification of MemMoRFs and demonstrate robust performance on transmembrane and membrane-associated proteins.
Abstract
Membrane molecular recognition features (MemMoRFs) are lipid-binding intrinsically disordered regions (IDRs) that undergo disorder-to-order transitions to mediate critical membrane dynamics. Consequently, their dysregulation is closely linked to severe human pathologies, including neurodegenerative diseases and viral infections. Despite their biological significance, annotations for MemMoRFs are scarce, limiting the accuracy of computational predictors. We introduce PreMemMoRF, a deep learning framework that leverages transfer learning to alleviate data scarcity. The model is pre-trained on linear interacting peptides (LIPs) with similar conformational transitions and fine-tuned on MemMoRF datasets, capturing generalizable binding-related sequence features. PreMemMoRF outperforms existing predictors across multiple metrics and demonstrates robust performance on transmembrane and membrane-associated proteins. It also performs consistently in short linear motif prediction, highlighting cross-task generalizability. Proteome-wide analysis in yeast shows that predicted scores exhibit systematic differences across distinct transmembrane topological regions and are consistent with established physicochemical constraints of membrane proteins. Collectively, these results validate PreMemMoRF as a robust and reliable computational framework for the large-scale identification of MemMoRFs.
Intrinsically disordered proteins and regions (IDPs/IDRs) mediate diverse cellular functions through binding segments whose functional properties are encoded in dynamic conformational ensembles rather than a single static state. Existing predictors of linear interacting peptides (LIPs) and molecular recognition features (MoRFs) rely primarily on sequence-derived features, leaving ensemble-level biophysical properties largely unexplored. Here, we introduce BindCORE, an ensemble-aware deep learning framework that integrates global, local, and pairwise biophysical descriptors to predict interaction sites within IDRs. These features are processed through a multi-scale architecture that enables information exchange between sequence- and ensemble-based global, local, and pairwise information. Across established LIP and MoRF benchmarks, BindCORE consistently improves performance over sequence-based baselines, demonstrating the predictive signals of ensemble-derived properties beyond sequence-based representations alone. Feature-attribution analyses reveal that pairwise descriptors are the dominant contributors to prediction, while solvent accessibility, backbone dihedral entropy, and global geometric properties provide complementary information. Feature-importance rankings vary substantially across ensemble flavours, indicating that different conformational generators encode distinct biophysical signatures of interaction-site propensity. Together, our results show that conformational ensembles contain interpretable determinants of LIP and MoRF binding residues and establish BindCORE as a general framework for incorporating biophysical information into the prediction of functional regions in intrinsically disordered proteins. BindCORE is freely available as a ready-to-use Google Colab notebook (BindCORE Colab notebook). Key Messages BindCORE integrates ensemble-derived biophysical descriptors to predict residue-level interaction sites in intrinsically disordered proteins. Ensemble-derived features improve prediction performance over state-of-the-art sequence-based methods on both LIP and MoRF benchmarks. Pairwise ensemble descriptors, especially contact and dynamic cross-correlation maps, provide the strongest signals for predicting interaction-site residues, while global chain geometry, solvent accessibility, and backbone dihedral preferences add complementary information.
Nicolas Buton, Luiz Felipe Piochi, Hammed Khakzad· bioRxiv· 0 citations
The results demonstrate the effectiveness of integrating multi-scale and multi-modal representations with cross-scale alignment for protein–RNA affinity prediction, and suggest that M2-PRNet can highlight relevant RNA-binding regions and support preliminary discrimination between strong and weak binders when plausible complex structures are available.
Junkai Wang, G. Luo, Yun-Song Yang et al.· Bioinformatics· 0 citations
Accurate prediction of variants within intrinsically disordered regions (IDRs) is crucial for advancing disease diagnosis and biomedical interpretation. However, the intrinsic lack of stable structural conformations and the high sequence variability of IDRs make it challenging for existing predictors to achieve robust performance in these regions. Here, we introduce DisoPatho, a deep learning framework specifically tailored for predicting disease-associated variants in IDRs. DisoPatho features a novel mutation-centric architecture that utilizes the variant site as an anchor for feature construction and interaction. The core innovation lies in a cross-view adaptive-feature interaction mechanism, which synergistically integrates IDR-specific energy representations with embeddings from protein language models, including xTrimoPGLM and Evolutionary Scale Modeling. This strategy enables the comprehensive capture of evolutionary constraints and physicochemical patterns without requiring explicit structural descriptors, multiple-sequence alignments, or hand-crafted conservation scores. Consequently, DisoPatho exhibits enhanced discriminative power better adapted to the highly flexible nature of IDRs. Comprehensive evaluations across multiple IDR data sets demonstrate that DisoPatho substantially outperforms existing methods. In 5-fold cross-validation, it achieves average AUCs of 0.899 and 0.840, with ACCs of 0.862 and 0.860 on two data sets. Notably, on a highly confounded independent test set where phylogenetic constraints offer limited discriminative signals, DisoPatho yields a 50.2% relative improvement in MCC over AlphaMissense on their respective predictable variants, while achieving broader prediction coverage. In-depth analyses of the prediction results further confirm the effectiveness and stability of the framework in IDR-specific scenarios. The code, data sets, and predictions for DisoPatho are available for academic use at https://github.com/IBHFLab/DisoPatho.
Xiaohua Wang, Shaojie Zhang, Hongmei Jiang et al.· Journal of Chemical Informat...· 0 citations
Identifying transcription factor binding sites (TFBSs) is fundamental to understanding complex gene regulatory mechanisms and the functions of non-coding regions. Although existing methods have achieved substantial strides, capturing both local structural features and long-range spatial dependencies within DNA sequences remains a major challenge for improving prediction accuracy. In this study, we propose DNCLA, a deep learning model that synergizes multisize convolutional fusion, Bidirectional Long ShortTerm Memory (Bi-LSTM) networks, and a multi-head self-attention
mechanism. At the feature extraction level, DNCLA breaks through the limitations of traditional single-sequence encoding by fusing Nucleotide Chemical Properties (NCP) with Dinucleotide Physicochemical Properties (DPCP). NCP provides a refined characterization of chemical differences between bases based on ring structures, hydrogen bond sites, and functional group properties, while DPCP introduces parameters such as local structural stability and geometric flexibility of the DNA. Subsequently, the model extracts spatial evolution from these high-dimensional features through a multi-size convolutional module; captures long-range spatial dependencies using Bi-LSTM layers; and employs a multi-head self-attention mechanism to achieve adaptive weight distribution of global features, thereby enhancing the perception of key regulatory motifs. Results from training and testing the proposed model on 165 ChIPseq datasets demonstrate that DNCLA possesses robust generalization capabilities and high predictive performance in TFBSs identification. This suggests that the incorporation of physicochemical features better elucidates the essence of interactions between transcription factors and DNA.
Jingjue Wei, Jie Feng· Match-communications in Math...· 0 citations
PLM-ArgMe is presented that is based on a symmetry-sensitive Transformer framework using context-aware ESM-2 residue embeddings, which is mapped through a novel Bio-Symmetric Mirrored Sinusoidal Encoding strategy to address the biological symmetry hypothesis of arginine methylation.
Nitika Bhatt, Kartik Joshi, R. Rout et al.· Biochemical and Biophysical...· 0 citations
DHST is proposed, a deep hybrid structure–topology framework that integrates sequence semantics from a pretrained protein language model with local structural information learned by a residual graph convolutional network and introduces site-specific persistent homology to encode multi-scale topological invariants and a topology-guided residue-wise gated fusion module to modulate structure–semantics representations using local topological embeddings.
Bin Lu, Fujun Xiang, Hai-Long Wang et al.· Applied Sciences· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.