MCGLPPI, a novel geometric representation learning framework that combines graph neural networks (GNNs) with the MARTINI molecular coarse-grained (CG) model to predict overall PPI properties accurately and efficiently, offers an effective and efficient solution for PPI overall property predictions.
Abstract
Structure-based machine learning algorithms have been utilized to predict the properties of protein-protein interaction (PPI) complexes, such as binding affinity, which is critical for understanding biological mechanisms and disease treatments. While most existing algorithms represent PPI complex graph structures at the atom-scale or residue-scale, these representations can be computationally expensive or may not sufficiently integrate finer chemical-plausible interaction details for improving predictions. Here, we introduce MCGLPPI, a novel geometric representation learning framework that combines graph neural networks (GNNs) with the MARTINI molecular coarse-grained (CG) model to predict overall PPI properties accurately and efficiently. This framework maps proteins onto a concise CG-scale complex graph, where nodes represent CG beads and edges encode chemically plausible interactions. The GNN-based encoder is tailored to extract high-quality representations from this graph, efficiently capturing the overall properties of the protein complex structure. Extensive experiments on three different downstream PPI property prediction tasks demonstrate that MCGLPPI achieves competitive performance compared with the counterparts at the atom- and residue-scale, but with only a third of the computational resource consumption. Furthermore, the CG-scale pre-training on protein domain-domain interaction structures enhances its predictive capabilities for PPI tasks. MCGLPPI offers an effective and efficient solution for PPI overall property predictions, serving as a promising tool for the large-scale analysis of biomolecular interactions.
This work introduces a scalable TDA framework for extracting information on PPIs directly from localized protein surface patches that leverages multiscale topological descriptors, evaluated from patch-wise point cloud representations of protein mesh surfaces, combined with supervised machine learning models for interface prediction.
Angan Mukherjee, ByungUk Park, Adam Malmstrom et al.· bioRxiv· 0 citations
An Algebraic Graph Neural Network model designed to encode molecular structures into a low-dimensional graph representation while preserving critical biochemical interactions is introduced, demonstrating superior performance in binding affinity prediction compared to state-of-the-art scoring functions.
Augustine Ouru, Xi Chen, Cameron Yeagle et al.· Computational and Mathematic...· 0 citations
This work presents HGRL-PPIS, a novel hierarchical graph representation learning approach for predicting protein-protein interaction sites that achieves superior performance over competing methods on multiple benchmark datasets, enabling more reliable detection of protein-protein binding residues.
An atom–bond bipartite graph modeling approach that treats atoms and bonds as explicit learnable node types and jointly models atom–atom, atom–bond, and bond–bond local interactions within a unified propagation framework is introduced.
Xing Zhao, Xianlai Chen, Yunbo Wang et al.· Bioinformatics· 0 citations
Accurate prediction of protein-protein interaction sites (PPISs) plays a crucial role in understanding protein function, elucidating disease mechanisms, and facilitating drug target discovery. Although conventional approaches based on sequence or structural features have shown promising results, they still face several challenges. These challenges include oversmoothing in deep graph neural networks (GNNs) and poor generalization to domain-specific data. To address these issues, we propose RGLLA-PPIS, a novel multimodal prediction model that integrates retrieval-augmented learning and residual GNNs for PPIS identification. In RGLLA-PPIS, protein graphs are constructed by combining AlphaFold3 (AF3)-predicted protein structures with multiple sequence-derived features. To effectively capture both local and global spatial dependencies, the model employs equivariant GNN (EGNN) and GCN modules with residual connections, which help alleviate the oversmoothing problem and preserve node-level variability. Moreover, during prediction, we used the retrieval-augmented knowledge provided by the pretrained protein language model (PLM) Evolla and ChatGPT-4o to construct semantic priors to supplement potential functional site information and enhance the generalization capacity of the prediction model. Extensive experiments on benchmark datasets show that RGLLA-PPIS outperforms several state-of-the-art baselines in both accuracy and robustness. Furthermore, comparison with wet-lab results on a domain-specific protein system reveals a strong correspondence between experimental functional sites and the high-probability regions predicted by RGLLA-PPIS. This demonstrates the model's potential to guide real-world protein engineering tasks. The source code can be found at: https://github.com/MiJia-ID/RGLLA-PPIS.
Jia Mi, Ya-Wen Liu, Chong Chu et al.· IEEE Transactions on Neural...· 0 citations
The prediction of molecular properties and biological functions across vastly different scales of macromolecules presents a formidable challenge in computational chemistry and bioinformatics. Traditional computational approaches often treat synthetic polymers and biological proteins as fundamentally distinct entities, applying specialized algorithms that fail to leverage the shared topological and chemical principles underlying both domains. This paper introduces a novel computational framework based on Transfer Message Passing, designed to unify the representation and predictive modeling of polymer and protein graphs. By conceptualizing both classes of macromolecules as complex, attributed graphs where nodes represent constituent functional units and edges denote chemical or spatial interactions, we establish a generalized topological space suitable for advanced graph neural networks. The proposed transfer learning mechanism dynamically adapts message passing operations learned from data-rich protein databases to infer complex physical and thermodynamic properties in specialized polymer datasets, mitigating the pervasive issue of data scarcity in polymer informatics. Extensive empirical evaluations demonstrate that our framework significantly outperforms domain-specific baseline models in predicting polymer bandgaps, glass transition temperatures, and protein enzymatic functions. Furthermore, detailed ablation studies reveal that the cross-domain attention mechanisms effectively align latent representations without compromising task-specific predictive accuracy. Ultimately, this research provides a robust theoretical foundation and a scalable computational tool for accelerated materials discovery and biomolecular engineering.
Richard Yat-Long Ma, Jessica Wing-Yan Lai· International Journal of Com...· 0 citations
Related blog posts
Microsoft Research Blog· microsoft.comJul 30, 2026
LLMs do not get smarter just by remembering more. EvoLib turns experience into evolving knowledge, taking reusable skills and insights that help models learn and adapt across tasks long after deployment. The post EvoLib: Turning experience into evolving knowledge appeared first on Microsoft Research.
What does it take to trust AI-driven HVAC optimization? Our AI Model Factory combines agents, machine learning, reinforcement learning and deterministic checks in a governed workflow designed for messy, real-world building data. The post We built an AI factory for HVAC control appeared first on GPT-Lab.
MIT News · Artificial Intelligence· news.mit.eduAug 27, 2026
A new machine-learning framework aims to improve the success rate of computational protein design while moving away from results that reproduce sequences found in nature.
MIT News · Artificial Intelligence· news.mit.eduAug 18, 2026
A new method for surgically removing training examples from a model reveals that as datasets grow, the link between what a model learns and what it produces dissolves.
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.