Jan 2026· Computational and Mathematical Biophysics· Vol 14· 0 citations· 55 references
TL;DR
An Algebraic Graph Neural Network model designed to encode molecular structures into a low-dimensional graph representation while preserving critical biochemical interactions is introduced, demonstrating superior performance in binding affinity prediction compared to state-of-the-art scoring functions.
Abstract
Abstract Predicting protein-ligand binding affinity is a fundamental challenge in drug discovery. Recent advances in deep learning have led to the development of numerous models, many of which rely on three-dimensional protein-ligand complex structures and focus primarily on affinity prediction. In this study, we introduce an Algebraic Graph Neural Network (AGNN) model designed to encode molecular structures into a low-dimensional graph representation while preserving critical biochemical interactions. While algebraic graph theory has been widely used in physical modeling and molecular studies, traditional methods often struggle to accurately capture the complexity of biomolecular interactions. To address this limitation, our proposed AGNN model leverages multiscale weighted colored subgraphs to describe molecular interactions through a graph neural network. These representations allow the model to effectively learn the geometric and topological features of protein-ligand complexes. The AGNN model integrates graph convolution layers and attention mechanisms to refine feature extraction and improve the interpretability of learned embeddings. Furthermore, we incorporate gradient boosting decision trees (GBDTs) to enhance the prediction of binding affinities by capturing nonlinear relationships between molecular features. Our approach is extensively validated using benchmark datasets, including PDBBind and CASF-2016, demonstrating superior performance in binding affinity prediction compared to state-of-the-art scoring functions.
Protein-ligand binding affinity (PLA) prediction aims to guide rational drug design by estimating the strength of interaction. The effectiveness of the representation learning of protein and ligand is key to successful PLA prediction. To this end, attention mechanism, as a powerful architectural paradigm, has been introduced and gradually emerged as the prevailing approach. However, intuitively, the classical attention paradigm based on similarity does not fit the biological mechanisms relevant for binding. Worse still, the cooperative and antagonistic effects among multiple atoms are deliberately disregarded in the classical formulation of attention mechanisms. Consequently, the rigid transplantation of classical architectures substantially undermines the PLA prediction performance. To address these challenges, we employ a hierarchical statistical attention model (HISA). Specifically, HISA employs a statistical attention mechanism (SAM) based on non-similarity computation to fit the biological prior and perceive the relationship of multiple atoms. In addition, we optimize HISA by employing clustering, enabling hierarchical representations of biomolecules. Extensive experiments demonstrate that HISA achieves state-of-the-art performance on multiple PLA benchmarks while simultaneously exhibiting generalizability and interpretability.
Changming Yao, Shunfanyi Li, Shanghui Deng et al.· IEEE transactions on computa...· 0 citations
Predicting drug–target interactions is critical for drug discovery, yet many deep learning methods overlook atom–residue–level relationships. We propose Protein Heterogeneous Graph learning for Drug–Target Interaction prediction (PHGDTI), a multimodal framework that integrates sequence and structural cues for binding prediction. Drug and protein sequences are embedded with Mol2Vec and Tasks Assessing Protein Embeddings (TAPE) and refined by a self-attention module. In parallel, a drug–protein graph encoder models three complementary graphs: a drug atom graph, a protein residue graph, and a heterogeneous atom–residue graph. Graph attention layers propagate intra- and intermolecular information, and SAGPooling yields compact structural representations. Fusing these structural and sequence features enables accurate affinity estimation. Experiments on the Davis kinase dataset and GalaxyDB dataset show PHGDTI surpasses competitive baselines, and ablation results highlight the benefit of heterogeneous graph modeling.
Hua Qian, Deng Pan, Liangpeng Nie et al.· Journal of Computational Bio...· 0 citations
HyBind-NN is developed, a multimodal graph neural network that integrates protein language models (PLMs) with 3D structural and dynamic datasets to predict protein–protein and protein–peptide affinity, and it is demonstrated that combining ESM-2 sequence embeddings with precise 3D Voronoi spatial geometry enables accurate affinity predictions across diverse structural datasets.
E. A. Bogdanova, A. Chernukhin, Alexey K. Shaytan· International Journal of Mol...· 0 citations
Despite challenges related to data sparsity and conformational variability, ViTs show strong performance and high robustness in structure-based affinity prediction tasks, underscore their effectiveness in learning spatial patterns and suggest broader applicability to related tasks, such as protein-protein or protein-nucleic acid interaction modeling.
Jakub Poziemski, Paweł Siedlecki· Scientific Reports· 0 citations
Deep neural network (DNN)-based in silico models show great promise in predicting the properties and bioactivities of novel compounds, including small molecules. Among traditional approaches, structure-based drug design (SBDD) remains a fundamental approach for drug discovery using molecular docking, scoring functions, and molecular dynamics simulations. However, these approaches are often constrained by limited flexibility, resolution, and generalizability. Geometric deep learning (GDL) offers a transformative alternative by enabling models to learn directly from non-Euclidean molecular representations, such as graphs, point clouds, and meshes, capturing critical 3D spatial relationships inherent to protein-ligand interactions. This review highlights the theoretical underpinnings and practical applications of GDL in small-molecule drug discovery, focusing on tasks including binding affinity prediction, virtual screening, de novo molecule generation, pose prediction, ADMET profiling, and protein flexibility modeling. We explore key GDL architectures, graph neural networks, SE(3)-equivariant networks, 3D convolutional neural networks, point cloud models, and geometric transformers, and assess their performance across various drug discovery benchmarks. The integration of geometry-aware AI models with experimental and computational workflows was also highlighted for its potential to streamline hit-to-lead optimization and advance rational drug design. Despite remarkable progress, the field faces challenges including limited high-quality 3D structural datasets, protein flexibility representation, and the interpretability of deep models. Addressing these issues through hybrid modeling approaches, multi-resolution learning, and self-supervised training could further elevate GDL's impact. Ultimately, GDL stands at the frontier of AI-enhanced pharmaceutical innovation, offering unprecedented precision, efficiency, and insight in the pursuit of next-generation therapeutics.
A. Srivastav, Unnati Modi, Rahul Kumar et al.· Molecular Informatics· 0 citations
Accurate identification of near-native ligand binding poses is a central challenge in structure-based drug design. From a physical point of view, the successful construction of a protein–ligand complex structure is dependent on whether protein and ligand can form enough atomically pairwise interactions that result in a global energy minimum. In this work, we report a machine learning scoring strategy for protein–ligand screening which explicitly considers the Native Contact Ratio (NCR), a topology inspired metric that quantifies the preservation of protein–ligand interfacial contacts as well as interaction energy. This physics-awared supervision strategy provides a simple but efficient gradient field that faithfully reflects the complicated protein energy landscape than conventional 3D coordinate-based objectives. Building on this principle, we present DeepNCR, an energy-informed Transformer framework that encodes approximate Coulombic and dispersive interaction potentials across the protein–ligand binding interface. Furthermore, we introduce a feature pruning step that compresses the interaction tensor from 1470 to 868 dimensions, further improving signal-to-noise ratio and directing model attention toward the interaction motifs critical for binding specificity. The model optimizes topological objectives and at inference drives pose refinement through a differentiable hybrid gradient field integrating predicted NCR and AutoDock Vina energetics. Extensive evaluation on the CASF-2016 benchmark and the 3D-DISCO cross-docking data set demonstrates consistently high performance: a Top-1 docking success rate of 94.7%, a 1% Enrichment Factor of 21.21 in virtual screening, and a Top-1 cross-docking success rate of 34.8%. Mechanistic analysis reveals that NCR-guided optimization enables decoy escaping from local energy minima and drives the recovery of disrupted native interactions, confirming that NCR captures the physical determinants of binding rather than mere geometric proximity.
Zhenqiang Zhang, Zhihao Wang, Yang Liu et al.· Journal of Chemical Informat...· 0 citations