This work introduces a scalable TDA framework for extracting information on PPIs directly from localized protein surface patches that leverages multiscale topological descriptors, evaluated from patch-wise point cloud representations of protein mesh surfaces, combined with supervised machine learning models for interface prediction.
Abstract
Protein-protein interactions (PPIs) govern a wide range of cellular functions. The ability to predict PPI interfaces from protein molecular surfaces is important for understanding protein function and enabling therapeutic discovery. While recent advances in structure-based learning, particularly molecular-surface geometric deep learning frameworks, have demonstrated that protein surfaces encode rich geometric and physicochemical information, such approaches often remain computationally intensive and data-hungry. Alternatively, topological data analysis (TDA) has emerged as a mathematically rigorous framework for extracting robust, multiscale shape information from complex data. In this work, we introduce a scalable TDA framework for extracting information on PPIs directly from localized protein surface patches. Our approach leverages multiscale topological descriptors, evaluated from patch-wise point cloud representations of protein mesh surfaces, combined with supervised machine learning models for interface prediction. On a full dataset of 3,362 proteins, the proposed approach substantially reduced computational cost relative to an established geometric deep learning method, MaSIF-site, decreasing preprocessing time from approximately 27 s/protein to 5-8 s/protein and total training time from approximately 6 h to 1-1.3 h. Importantly, this computational reduction is achieved while maintaining mean test area under the receiver operating characteristic curve (AUC) values of 0.76 and 0.77 for patch radii of 9 Å and 12 Å, respectively, thus approaching the MaSIF-site test AUC of 0.84. Our results suggest that topology offers a scalable and computationally efficient approach for high-throughput extraction of information from complex biomolecular interfaces.
This work presents HGRL-PPIS, a novel hierarchical graph representation learning approach for predicting protein-protein interaction sites that achieves superior performance over competing methods on multiple benchmark datasets, enabling more reliable detection of protein-protein binding residues.
DHST is proposed, a deep hybrid structure–topology framework that integrates sequence semantics from a pretrained protein language model with local structural information learned by a residual graph convolutional network and introduces site-specific persistent homology to encode multi-scale topological invariants and a topology-guided residue-wise gated fusion module to modulate structure–semantics representations using local topological embeddings.
Bin Lu, Fujun Xiang, Hai-Long Wang et al.· Applied Sciences· 0 citations
Accurate identification of near-native ligand binding poses is a central challenge in structure-based drug design. From a physical point of view, the successful construction of a protein-ligand complex structure is dependent on whether protein and ligand can form enough atomically pairwise interactions that result in a global energy minimum. In this work, we report a machine learning scoring strategy for protein-ligand screening which explicitly considers the Native Contact Ratio (NCR), a topology inspired metric that quantifies the preservation of protein-ligand interfacial contacts as well as interaction energy. This physics-awared supervision strategy provides a simple but efficient gradient field that faithfully reflects the complicated protein energy landscape than conventional 3D coordinate-based objectives. Building on this principle, we present DeepNCR, an energy-informed Transformer framework that encodes approximate Coulombic and dispersive interaction potentials across the protein-ligand binding interface. Furthermore, we introduce a feature pruning step that compresses the interaction tensor from 1470 to 868 dimensions, further improving signal-to-noise ratio and directing model attention toward the interaction motifs critical for binding specificity. The model optimizes topological objectives and at inference drives pose refinement through a differentiable hybrid gradient field integrating predicted NCR and AutoDock Vina energetics. Extensive evaluation on the CASF-2016 benchmark and the 3D-DISCO cross-docking data set demonstrates consistently high performance: a Top-1 docking success rate of 94.7%, a 1% Enrichment Factor of 21.21 in virtual screening, and a Top-1 cross-docking success rate of 34.8%. Mechanistic analysis reveals that NCR-guided optimization enables decoy escaping from local energy minima and drives the recovery of disrupted native interactions, confirming that NCR captures the physical determinants of binding rather than mere geometric proximity.
Zhen-Qiang Zhang, Zhihao Wang, Yang Liu et al.· Journal of Chemical Informat...· 0 citations
PML is introduced, a novel computational framework that describes a binding interface as a family of multiscale manifolds, and results indicate that much of what determines binding strength is encoded in the shape of the interface itself, and that a single geometric description serves both classes without hand-tailored features.
Xingjian Xu, Zhe Su, Guo-Wei Wei et al.· arXiv.org· 0 citations
Predicting protein–protein binding free energy (ΔG) from structure remains a central challenge in computational biophysics. Here, we present GULP (Graph-based Unified Learning for Protein binding), a graph neural network (GNN) that jointly learns from a residue-level graph representation of the binding interface and global physicochemical descriptors. We systematically investigate how training data distribution affects model performance by comparing a full training set with a balanced subset enriched for extreme-affinity complexes. GULP is computationally efficient and provides interpretable insights into residue-level and physicochemical contributions to binding. On external validation, GULP achieves a mean absolute error (MAE) of 2.31 kcal/mol and shows moderate agreement with experimental ΔG values (Pearson r = 0.54, Spearman ρ = 0.58).
The prediction of molecular properties and biological functions across vastly different scales of macromolecules presents a formidable challenge in computational chemistry and bioinformatics. Traditional computational approaches often treat synthetic polymers and biological proteins as fundamentally distinct entities, applying specialized algorithms that fail to leverage the shared topological and chemical principles underlying both domains. This paper introduces a novel computational framework based on Transfer Message Passing, designed to unify the representation and predictive modeling of polymer and protein graphs. By conceptualizing both classes of macromolecules as complex, attributed graphs where nodes represent constituent functional units and edges denote chemical or spatial interactions, we establish a generalized topological space suitable for advanced graph neural networks. The proposed transfer learning mechanism dynamically adapts message passing operations learned from data-rich protein databases to infer complex physical and thermodynamic properties in specialized polymer datasets, mitigating the pervasive issue of data scarcity in polymer informatics. Extensive empirical evaluations demonstrate that our framework significantly outperforms domain-specific baseline models in predicting polymer bandgaps, glass transition temperatures, and protein enzymatic functions. Furthermore, detailed ablation studies reveal that the cross-domain attention mechanisms effectively align latent representations without compromising task-specific predictive accuracy. Ultimately, this research provides a robust theoretical foundation and a scalable computational tool for accelerated materials discovery and biomolecular engineering.
Richard Yat-Long Ma, Jessica Wing-Yan Lai· International Journal of Com...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.