Aug 2026· Bioinformatics· Vol 42· 0 citations· 36 references
Medicine
TL;DR
A structure-based method using SE(3)-transformers to learn residue compatibility with the local structural environment from experimentally resolved kinase 3D structures that captures biologically meaningful relationships between residue identity and 3D structural context is presented.
Abstract
Abstract Motivation Protein kinases are key regulators of cellular signaling and are frequently implicated in human diseases. Although kinase domains are structurally conserved, predicting the effects of amino acid substitutions remains challenging as mutations often introduce subtle structural perturbations that are not captured by sequence-based or evolutionary methods. Existing supervised approaches further rely on pathogenicity annotations that are inconsistent across databases, thereby motivating the development of structure-based, label-independent frameworks for mutation effect prediction. Results We present a structure-based method using SE(3)-transformers to learn residue compatibility with the local structural environment from experimentally resolved kinase 3D structures. Proteins are represented as atom-level graphs with physicochemical descriptors derived from the CHARMM force field and spatial connectivity. The model is trained on two self-supervised tasks given local structural context: masked residue atom reconstruction and masked residue classification. This formulation enables learning of geometric and physicochemical constraints without relying on pathogenicity labels. Evaluation using reconstruction loss, residue prediction accuracy, and comparison with BLOSUM substitution patterns indicate that the model captures biologically meaningful relationships between residue identity and 3D structural context. We interpret the scores assigned to alternative amino acids as measures of structural fitness, where low-scoring residues are hypothesized to be less compatible with the local environment and more likely to induce deleterious effects on protein structure and activity. Availability and implementation https://zenodo.org/records/20393799.
DHST is proposed, a deep hybrid structure–topology framework that integrates sequence semantics from a pretrained protein language model with local structural information learned by a residual graph convolutional network and introduces site-specific persistent homology to encode multi-scale topological invariants and a topology-guided residue-wise gated fusion module to modulate structure–semantics representations using local topological embeddings.
Bin Lu, Fujun Xiang, Hai-Long Wang et al.· Applied Sciences· 0 citations
The interaction between proteins and metal ions plays a pivotal role in the pathogenesis of various diseases. Base pair substitution subsequently triggers single-amino acid substitutions, leading to missense mutations that may impair protein function and consequently cause severe human diseases. Predicting the pathogenicity or benign nature of mutation sites in metalloproteins is crucial for early disease diagnosis and the acceleration of innovative drug development. In this study, we propose MetalDiagnosis, a deep learning framework that integrates an improved equivariant graph neural network with a pre-trained protein language model to predict disease-associated mutation sites in metal-binding proteins. Specifically, MetalDiagnosis employs a sliding-window strategy to extract deep contextual semantic information from the pre-trained protein language model ESM Cambrian, thereby overcoming the limitations of manual feature extraction. In parallel, three-dimensional geometric features are extracted by the equivariant graph neural network, which learns global patterns through positional encoding and virtual nodes. MetalDiagnosis integrates sequence and structural representations at both the residue and geometric levels to ensure comprehensive feature incorporation. Experimental results demonstrate that MetalDiagnosis outperforms state-of-the-art methods on the independent test set. When applied to 611 variants of uncertain significance (VUS), it produces fewer ambiguous predictions and reclassifies more VUS into likely pathogenic or likely benign categories with higher confidence. Furthermore, case studies demonstrate that MetalDiagnosis can accurately identify pathogenic mutations located in functionally critical regions. These results suggest that MetalDiagnosis offers an effective computational framework for identifying disease-associated mutations in metalloproteins.
Xudong Guo, Runchang Jia, Heyun Sun et al.· IEEE journal of biomedical a...· 0 citations
LINKER is the first sequence-based model to predict residue-functional group interactions according to biologically defined interaction types, using only a protein sequence and the SMILES representation of the ligand, and requires only sequence-level input at inference.
Phuc Pham, Viet Thanh Duy Nguyen, Kevin Song et al.· Journal of Chemical Informat...· 1 citation
UniStab is introduced, an end-to-end framework for predicting stability changes across all mutation types by leveraging the implicit geometric reasoning of a pre-trained folding model and demonstrates state-of-the-art performance, particularly in the challenging scenarios of multi-point mutations and indels.
Hong Tan, Sheng-Geng Lin, Yi Xiong· Chemical Science· 0 citations
By combining protein language model embeddings with topology-adaptive geometric reasoning, DiConSite offers a reusable framework for residue-level protein interaction analysis and achieves consistently strong and often best-performing results, while improving robustness to structural uncertainty and cross-modal variation.
Shou-Zhi Chen, Zhenchao Tang, Linlin You et al.· IEEE Transactions on Pattern...· 1 citation
Intrinsically disordered proteins and regions (IDPs/IDRs) mediate diverse cellular functions through binding segments whose functional properties are encoded in dynamic conformational ensembles rather than a single static state. Existing predictors of linear interacting peptides (LIPs) and molecular recognition features (MoRFs) rely primarily on sequence-derived features, leaving ensemble-level biophysical properties largely unexplored. Here, we introduce BindCORE, an ensemble-aware deep learning framework that integrates global, local, and pairwise biophysical descriptors to predict interaction sites within IDRs. These features are processed through a multi-scale architecture that enables information exchange between sequence- and ensemble-based global, local, and pairwise information. Across established LIP and MoRF benchmarks, BindCORE consistently improves performance over sequence-based baselines, demonstrating the predictive signals of ensemble-derived properties beyond sequence-based representations alone. Feature-attribution analyses reveal that pairwise descriptors are the dominant contributors to prediction, while solvent accessibility, backbone dihedral entropy, and global geometric properties provide complementary information. Feature-importance rankings vary substantially across ensemble flavours, indicating that different conformational generators encode distinct biophysical signatures of interaction-site propensity. Together, our results show that conformational ensembles contain interpretable determinants of LIP and MoRF binding residues and establish BindCORE as a general framework for incorporating biophysical information into the prediction of functional regions in intrinsically disordered proteins. BindCORE is freely available as a ready-to-use Google Colab notebook (BindCORE Colab notebook). Key Messages BindCORE integrates ensemble-derived biophysical descriptors to predict residue-level interaction sites in intrinsically disordered proteins. Ensemble-derived features improve prediction performance over state-of-the-art sequence-based methods on both LIP and MoRF benchmarks. Pairwise ensemble descriptors, especially contact and dynamic cross-correlation maps, provide the strongest signals for predicting interaction-site residues, while global chain geometry, solvent accessibility, and backbone dihedral preferences add complementary information.
Nicolas Buton, Luiz Felipe Piochi, Hammed Khakzad· bioRxiv· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.