BiteNetI is a structure-based deep learning model that uses 3D convolutional neural networks to simultaneously localize ion-binding centers and predict binding residues for 14 biologically relevant ions, supporting comprehensive and large-scale annotation of protein-ion interactions.
Abstract
Protein-ion interactions are essential for many cellular processes, including enzymatic catalysis, signaling, and allosteric regulation. However, mapping ion-binding sites experimentally remains labor-intensive and expensive. Here, we present BiteNetI, a structure-based deep learning model that uses 3D convolutional neural networks to simultaneously localize ion-binding centers and predict binding residues for 14 biologically relevant ions. Trained on a carefully curated dataset of over 10,000 high-resolution protein–ion complexes, in which near-identical binding sites are consistently annotated by transferring ions between homologous structures, BiteNetI shows strong generalization ability across diverse ions within a unified multitask architecture. On two different test benchmarks, BiteNetI achieves state-of-the-art performance compared to existing ion-binding predictors as well as to a more general method, AlphaFold3, when used to predict the entire structure of protein bound to ions. Finally, for physiologically relevant ions such as Ca2+, Na+ and K+, BiteNetI achieves two- to three-fold improvement in accuracy. Structure-based deep learning enables accurate prediction of protein-ion binding sites across 14 biologically relevant ions, supporting comprehensive and large-scale annotation of protein-ion interactions.
A state-space–driven deep learning framework that leverages the efficient long-range modeling capability of the Vision Mamba architecture to learn from three-dimensional protein surfaces represented as two-dimensional geometric or physicochemical grids, establishing state-space models as efficient, interpretable, and scalable architectures for molecular surface learning.
Metal ions serve as essential cofactors in approximately 30%–40% of proteins, and accurate recognition of their binding sites is central to function annotation, drug discovery, and metalloenzyme design. Existing predictors often operate at residue level, generate many false positives, or depend strongly on high-quality bound structures. We present DeepMetal, a hierarchical coarse-to-fine framework that combines ESM-2 residue screening, biophysics-constrained Dynamic Center-Iterative Clustering (DCIC), and a site-level SE(3)-equivariant graph neural network for candidate-site validation and metal typing. On a non-redundant BioLiP2-derived benchmark, DeepMetal achieves an AUROC of 0.775 and an F2 score of 0.533 for transition-metal site localization, outperforming representative baselines MetalNet2 and PinMyMetal under the same intersectional evaluation setting. These results show that sequence-driven screening, geometry-aware assembly, and equivariant validation can jointly improve practical metal-binding site prediction from predicted protein structures.
Bing Liu, Yangfan Xu, Yunpeng Wang et al.· ACM International Conference...· 0 citations
Predicting the location of metal-binding sites in proteins is crucial for fundamental biological questions and biotechnological applications. Over the past decade, the rise in metal-bound protein structures in the Protein Data Bank, combined with advanced statistical models such as deep learning, has accelerated the development of metal-binding site prediction tools. Several approaches are now available, offering high-quality benchmarks and predictive performance. Our initial development in this area is BioMetAll, whose first version was based on backbone pre-organization. Here, we introduce its second version, featuring two major updates: 1) metal-specific scoring functions and 2) prediction using backbone geometry alone or in combination with first coordination sphere descriptors. Apart from demonstrating metal sensitivity and yielding better benchmarking results, this new version allows the assessment of the influence of considering the metal’s first coordination sphere versus backbone pre-organization on how metallic species bind to proteins.
Protein phosphorylation regulates signaling, yet atomic-level substrate specificity remains elusive due to sparse structural data and phosphorylation-site-insensitive deep-learning predictors. Here we present a pipeline reformulating kinase-substrate modeling as a Bayesian inference problem. By integrating curated data sets and literature evidence parsed by Large Language Models, we converted diverse biological knowledge into structural restraints for the restraint-guided deep-learning model GRASP. For EGFR, BRAF and JNK1, we obtained 336 new phosphorylation-site-specific structure candidates refined by molecular dynamics. These models recapitulate known features, such as JNK1's hydrophobic docking groove, and enabled a Virtual Position Scanning Peptide Array (V-PSPA) to map recognition patches and derive sequence preferences. Cross-referencing predicted interfaces with AlphaMissense pathogenicity scores reveal that the interaction types and distances to the catalytic pocket significantly influence pathogenicity scores. A comparison with clinical mutation data sets further connects pathogenic mutations to the kinase-substrate interface. This high-resolution, high-throughput pipeline can be broadly applicable to kinase specificity studies and general drug discovery.
Jinyuan Hu, Shimian Li, Yue Xue et al.· Journal of Chemical Informat...· 0 citations
The Ca2+ binding sites of proteins are critical for their function, particularly in processes such as signal transduction, enzyme regulation, and structural stability. In this study, the calcium-binding sites of NtEhCaBP1 (Entamoeba histolytica calcium-binding protein). This paper proposes Statistical Ranking Deep Learning (SR-ML) to estimate the binding affinities of ten protein variants, The proposed SR-ML model computes the features in the proteins with the detection of sequences in the bindings. The classification of binding sites evaluated with the optimization of the features. With each predicted variant’s binding affinity correlates well with its experimental value with Kendall Tau (τ) values ranging from 0.78 to 0.95 and Spearman rank correlation (ρ) ranging from 0.75 to 0.94. Specifically, the Root Mean Square, Deviation (RMSD) shows protein flexibility in values of 0.95 to 1.50 angstrom and Root Mean Fluctuation (RMSF) values of 0.30 angstrom to 0.50 angstrom. The binding energy falls from negative 4.90 kcal/mol to negative 7.20 kcal/mol proposing differing levels of protein stability. Secondly, considering calcium coordination geometry we describe how there are octahedral, tetrahedral and trigonal bipyramidal structures in various proteins, with Kd values of 0.3 uM to 5.0 uM. The anti-AIDS bioactive example of mutagenesis validation is at a 120-folds to 600-folds increase from binding affinity for several mutations involving dynamic correlation with values of between 0.88 to 0.97. These outcomes reveal that the SR-ML model has certain predictive preciseness in terms of the Ca-binding sites and protein motions, which is valuable for Drug designing involving the Ca signalling Pathway.
P. Parwekar, S. Gourinath, Jaishree Jain et al.· PLoS ONE· 0 citations
OrgNet+, a conformational ensemble-aware and orientation-gnostic framework that explicitly incorporates protein structure flexibility during training, is introduced, which substantially reduces intra-ensemble prediction variance while simultaneously improving predictive accuracy.
A. Sarycheva, Aleksandr Shumilov, Petr Popov· Bioinformatics· 0 citations