An interpretable ensemble machine-learning framework that integrates protein language model embeddings of sequence and predicted structure to predict LLPS propensity and classify proteins as self-separating or partner-dependent and identifies critical phase-separating regions and quantifies mutation-induced perturbations in LLPS.
Abstract
Coding mutations within intrinsically disordered regions (IDRs) of proteins are increasingly implicated in human diseases yet remain poorly interpreted by conventional variant-effect predictors that rely on structural stability and conservation-based metrics. Quantifying disruption of IDR-mediated liquid-liquid phase separation (LLPS) offers a biophysically principled approach to interpreting the pathogenic impact of such variants. However, existing LLPS predictors suffer from training biases toward self-separating proteins, show limited performance on partner- dependent phase separation, and often lack interpretability for variant prioritization. We present an interpretable ensemble machine-learning framework that integrates protein language model embeddings of sequence and predicted structure to predict LLPS propensity and classify proteins as self-separating or partner-dependent. Our two-step classifiers outperform existing methods on independent benchmark datasets, with the largest gains for partner-dependent LLPS proteins. Beyond classification, our framework identifies critical phase-separating regions and quantifies mutation-induced perturbations in LLPS. Applied to disease-associated variant databases, we found that pathogenic mutations are enriched in predicted phase-separating regions and frequently perturb LLPS propensity scores, implicating mutation-induced LLPS dysregulation as a potential pathogenic mechanism for numerous diseases. Overall, our framework provides an accurate, interpretable approach for identifying phase-separating proteins and linking aberrant phase- separation behavior to disease pathogenesis.
The discordance persists: pLDDT correlates positively with PUNCH2 disorder in random and de novo proteins and negatively with β-strand fraction, opposite to the conserved and disordered baselines, a concrete failure mode that protein designers and other working on sequences remote in sequence space should be aware of when relying on predictor outputs.
Lars A. Eicholt, Lasse Middendorf· bioRxiv· 0 citations
Proteome-wide prediction and structural modeling of disordered protein interaction interfaces advance characterization of disease-associated variants in disordered protein regions.
D. Hubrich, Jesús Alvarado Valverde, C. Y. Lee et al.· Nature Structural & Molecula...· 0 citations
Accurate prediction of variants within intrinsically disordered regions (IDRs) is crucial for advancing disease diagnosis and biomedical interpretation. However, the intrinsic lack of stable structural conformations and the high sequence variability of IDRs make it challenging for existing predictors to achieve robust performance in these regions. Here, we introduce DisoPatho, a deep learning framework specifically tailored for predicting disease-associated variants in IDRs. DisoPatho features a novel mutation-centric architecture that utilizes the variant site as an anchor for feature construction and interaction. The core innovation lies in a cross-view adaptive-feature interaction mechanism, which synergistically integrates IDR-specific energy representations with embeddings from protein language models, including xTrimoPGLM and Evolutionary Scale Modeling. This strategy enables the comprehensive capture of evolutionary constraints and physicochemical patterns without requiring explicit structural descriptors, multiple-sequence alignments, or hand-crafted conservation scores. Consequently, DisoPatho exhibits enhanced discriminative power better adapted to the highly flexible nature of IDRs. Comprehensive evaluations across multiple IDR data sets demonstrate that DisoPatho substantially outperforms existing methods. In 5-fold cross-validation, it achieves average AUCs of 0.899 and 0.840, with ACCs of 0.862 and 0.860 on two data sets. Notably, on a highly confounded independent test set where phylogenetic constraints offer limited discriminative signals, DisoPatho yields a 50.2% relative improvement in MCC over AlphaMissense on their respective predictable variants, while achieving broader prediction coverage. In-depth analyses of the prediction results further confirm the effectiveness and stability of the framework in IDR-specific scenarios. The code, data sets, and predictions for DisoPatho are available for academic use at https://github.com/IBHFLab/DisoPatho.
Xiaohua Wang, Shaojie Zhang, Hongmei Jiang et al.· Journal of Chemical Informat...· 0 citations
Three modeling frameworks are developed, including models based on handcrafted features, models using embedding representations extracted from ProteinMPNN, and ensemble models integrating a diverse set of state‐of‐the‐art predictors integrating a diverse set of state‐of‐the‐art predictors.
Yang Liu, Jian Zhang, Minghui Li· Protein Science· 0 citations
Intrinsically disordered proteins and regions (IDPs/IDRs) mediate diverse cellular functions through binding segments whose functional properties are encoded in dynamic conformational ensembles rather than a single static state. Existing predictors of linear interacting peptides (LIPs) and molecular recognition features (MoRFs) rely primarily on sequence-derived features, leaving ensemble-level biophysical properties largely unexplored. Here, we introduce BindCORE, an ensemble-aware deep learning framework that integrates global, local, and pairwise biophysical descriptors to predict interaction sites within IDRs. These features are processed through a multi-scale architecture that enables information exchange between sequence- and ensemble-based global, local, and pairwise information. Across established LIP and MoRF benchmarks, BindCORE consistently improves performance over sequence-based baselines, demonstrating the predictive signals of ensemble-derived properties beyond sequence-based representations alone. Feature-attribution analyses reveal that pairwise descriptors are the dominant contributors to prediction, while solvent accessibility, backbone dihedral entropy, and global geometric properties provide complementary information. Feature-importance rankings vary substantially across ensemble flavours, indicating that different conformational generators encode distinct biophysical signatures of interaction-site propensity. Together, our results show that conformational ensembles contain interpretable determinants of LIP and MoRF binding residues and establish BindCORE as a general framework for incorporating biophysical information into the prediction of functional regions in intrinsically disordered proteins. BindCORE is freely available as a ready-to-use Google Colab notebook (BindCORE Colab notebook). Key Messages BindCORE integrates ensemble-derived biophysical descriptors to predict residue-level interaction sites in intrinsically disordered proteins. Ensemble-derived features improve prediction performance over state-of-the-art sequence-based methods on both LIP and MoRF benchmarks. Pairwise ensemble descriptors, especially contact and dynamic cross-correlation maps, provide the strongest signals for predicting interaction-site residues, while global chain geometry, solvent accessibility, and backbone dihedral preferences add complementary information.
Nicolas Buton, Luiz Felipe Piochi, Hammed Khakzad· bioRxiv· 0 citations
A structure-based method using SE(3)-transformers to learn residue compatibility with the local structural environment from experimentally resolved kinase 3D structures that captures biologically meaningful relationships between residue identity and 3D structural context is presented.
Shakiba Fadaei, F. Krebs, V. Zoete· Bioinformatics· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.