Skip to content
Open access

BiMba: using Vision Mamba to predict protein sites that bind other proteins

Jul 2026 · Bioinformatics · Vol 42 · 0 citations · 44 references
Medicine

TL;DR

A state-space–driven deep learning framework that leverages the efficient long-range modeling capability of the Vision Mamba architecture to learn from three-dimensional protein surfaces represented as two-dimensional geometric or physicochemical grids, establishing state-space models as efficient, interpretable, and scalable architectures for molecular surface learning.

Abstract

Abstract Motivation Identifying protein binding sites in protein–protein complexes is a central challenge in structural biology. Binding sites, consisting of groups of residues, govern how proteins recognize, and interact with protein partners. Thus, identifying them is essential for understanding biological function and guiding the design of effective biomolecules and even drug molecules. Despite major progress in computational approaches, their performance remains limited because most models underrepresent the combined influence of surface properties and residue-level information, leaving room for improvement. Recent advances in state-space models and vision-based deep learning offer an opportunity to address these limitations by efficiently modeling long-range spatial dependencies on protein surfaces. Here, we introduce BiMba (protein Binding site prediction using Vision Mamba), a state-space–driven deep learning framework that leverages the efficient long-range modeling capability of the Vision Mamba architecture to learn from three-dimensional (3D) protein surfaces represented as two-dimensional (2D) geometric or physicochemical grids. Results BiMba integrates complementary sources of information, capturing geometric and physicochemical determinants of molecular recognition as surface patches, encoded as 2D images, along with residue-level descriptors, yielding a unified representation that couples spatial topology with biochemical context. BiMba demonstrates competitive performance across diverse and specialized benchmark datasets, often outperforming existing state-of-the-art methods. In addition, BiMba incorporates perturbation-based and gradient-based interpretability analyses by extracting hidden attentions from Mamba layers, enabling visualization of feature relevance and biologically meaningful residue clusters. Overall, our findings establish state-space models as efficient, interpretable, and scalable architectures for molecular surface learning, advancing the application of deep learning in structural bioinformatics. Availability and implementation The BiMba source code, training, test, and benchmark datasets are available at https://github.com/Azam-Shi/BiMba.

Read PDF

Similar papers

Open access Jul 2026

BioMetAll v2.0: Introducing Scores, Metal Discrimination, and Side-Chain Descriptors for Predicting Metal-Binding Sites in Proteins

Predicting the location of metal-binding sites in proteins is crucial for fundamental biological questions and biotechnological applications. Over the past decade, the rise in metal-bound protein structures in the Protein Data Bank, combined with advanced statistical models such as deep learning, has accelerated the development of metal-binding site prediction tools. Several approaches are now available, offering high-quality benchmarks and predictive performance. Our initial development in this area is BioMetAll, whose first version was based on backbone pre-organization. Here, we introduce its second version, featuring two major updates: 1) metal-specific scoring functions and 2) prediction using backbone geometry alone or in combination with first coordination sphere descriptors. Apart from demonstrating metal sensitivity and yielding better benchmarking results, this new version allows the assessment of the influence of considering the metal’s first coordination sphere versus backbone pre-organization on how metallic species bind to proteins.

José-Emilio Sánchez-Aparicio, Raúl Fernández Díaz, Raúl Peña Losada et al. · 0 citations
Open access Sep 2025

LINKER: Learning Interactions between Functional Groups and Residues with Chemical Knowledge‑Enhanced Reasoning and Explainability

Accurate identification of interactions between protein residues and ligand functional groups is critical for understanding molecular recognition and guiding rational drug design. Existing deep learning approaches for protein–ligand interpretability typically rely on three-dimensional structural input or distance-based contact labels, which limit both their applicability and biological relevance. Here, we present LINKER, the first sequence-based model to predict residue-functional group interactions according to biologically defined interaction types, using only a protein sequence and the SMILES representation of the ligand. LINKER is trained via structure-supervised interaction learning, in which interaction labels are derived from three-dimensional protein–ligand complexes through functional group-based motif extraction. By representing ligands as ensembles of functional groups, the model emphasizes chemically meaningful substructures rather than mere spatial proximity. Importantly, LINKER requires only sequence-level input at inference, enabling large-scale applications in contexts where structural data are unavailable. Extensive experiments demonstrate that LINKER consistently outperforms established baselines, highlighting the utility of functional group abstractions and structure-based supervision for interpretable protein–ligand interaction prediction. Our source code is publicly available at: https://github.com/HySonLab/LINKER/.

Phuc Pham, Viet Thanh Duy Nguyen, Truong-Son Hy · 1 citation
Open access 2026

A Biologically Informed Hybrid Stacking Framework for Protein–Protein Interaction Prediction

Mapping the protein interactome is fundamental to understanding disease mechanisms and facilitating therapeutic development. Although protein language models (PLMs) such as ESM-2 have advanced protein-protein interaction (PPI) prediction, their high-dimensional representations remain difficult to connect to verifiable biological signals. To address this limitation, we propose HybridStack-PPI, a gray-box framework that combines ESM-2 sequence representations with explicit physicochemical and motif-derived biological descriptors. The architecture uses motif-anchored local pooling global mean pooling, symmetric pair encoding, fold-internal feature selection, LightGBM branch learners, and an elastic-net logistic-regression stacking layer. We evaluated the method using a C3 cluster-based cross-validation protocol with a 40% sequence-identity clustering threshold and a Same-GO hard-negative setting in which negative candidates shared functional annotations with positive pairs. Under this setting, HybridStack-PPI reached a Human ROC-AUC of 73.65%, PR-AUC of 91.35%, MCC of 28.06%, and specificity of 75.61%. The results indicate a conservative operating point: compared to more recall-oriented baselines, the proposed stack trades lower recall and F1 for higher specificity, MCC, and ranking behavior under functionally similar negative samples. We further reported cross-species transfer, ablation, latency, SHAP-based descriptor attribution, and meta-learner coefficient analyses to clarify both the promise and limitations of biologically informed PPI prediction.

T. T. Nguyen, X. Mai, N. Nguyen · 0 citations
Open access Jul 2026

Multivalent ion binding site identification with structure-based deep learning

BiteNetI is a structure-based deep learning model that uses 3D convolutional neural networks to simultaneously localize ion-binding centers and predict binding residues for 14 biologically relevant ions, supporting comprehensive and large-scale annotation of protein-ion interactions.

Igor Kozlovskii, Petr Popov · 0 citations
Open access Jun 2026

Pep2Mol: 3D Molecule Generation Targeting Protein-Protein Interfaces with Diffusion Models

P Pep2Mol is introduced, a diffusion-based generative model for 3D molecule design that targets orthosteric PPI sites by explicitly incorporating binding peptides or proteins as structural guidance, moving beyond conventional pocket-conditioned generation.

Rongting Yue, Zekun Yang, G. Seabra et al. · 0 citations
Open access Jun 2026

Hybrid Approach to Protein–Protein Complex Affinity Prediction Based on Language Models and Molecular Dynamics

HyBind-NN is developed, a multimodal graph neural network that integrates protein language models (PLMs) with 3D structural and dynamic datasets to predict protein–protein and protein–peptide affinity, and it is demonstrated that combining ESM-2 sequence embeddings with precise 3D Voronoi spatial geometry enables accurate affinity predictions across diverse structural datasets.

E. A. Bogdanova, A. Chernukhin, Alexey K. Shaytan · 0 citations