Skip to content
Open access

Predicting membrane protein localization by deep learning on structure and chemistry

Aug 2026 · Protein Science · Vol 35 · 0 citations · 47 references
Medicine

TL;DR

A graph neural network model of proteins was trained on experimentally determined membrane protein structures to predict the native membrane environment of transmembrane domains from their structure, and the algorithm, “GPSforTMDs,” obtains overall performance that is competitive with sequence‐based methods.

Abstract

It has been known since at least the 1980's that the structure and chemistry of membranes and membrane proteins are matched. Exploiting this fact, a graph neural network model of proteins was trained on experimentally determined membrane protein structures to predict the native membrane environment of transmembrane domains from their structure. The algorithm, “GPSforTMDs,” learns to generalize about membrane protein structure, obtains overall performance that is competitive with sequence‐based methods, and obtains exceptional performance for some categories of membrane environment, even when training examples are few. Other categories it finds more challenging, in some cases for clear reasons (for example, compatibility of TMDs with membranes along the secretory pathway), and in other cases that are mysterious (mistaking archaeal TMDs for bacterial, and vice versa). The results motivate the need for high quality databases reporting TMD localization, and suggest that peering inside the algorithm will reveal new “rules” for membrane proteins. The code and associated database is available at https://github.com/bivekpok/GPSforTMDs.

Read PDF

Similar papers

Open access Aug 2026

DeepTMHMM2 enables accurate prediction of transmembrane protein topology and subcellular location

Transmembrane α-helical and β-barrel proteins are a ubiquitous component of proteomes. Topology prediction infers how proteins are embedded in lipid bilayers, identifying membrane-spanning segments and their orientation. While recent methods achieve high performance for membrane-spanning segments, they cannot predict re-entrant regions and interfacial helices – membrane-associated segments that partially insert but do not cross the bilayer – nor identify which biological membrane a protein resides in. Here, we present DeepTMHMM2, the first predictor to include re-entrant regions and interfacial helices in its topologies and jointly predict localization across 17 biological membranes. Benchmark results show that DeepTMHMM2 successfully learns to predict the additional elements, while achieving strong performance on canonical α-helical and β-barrel topology prediction. Applying DeepTMHMM2 to Swiss-Prot reveals that non-crossing segments are a ubiquitous feature of the transmembrane proteome, with interfacial helices present in nearly a quarter of all α-helical transmembrane proteins.

Felix Teufel, Jeppe Hallgren, Henrik Nielsen et al. · 0 citations
Open access Jul 2026

Predictions of protein–protein interactions: Learning sequences and structures

A neural network-based pipeline that integrates amino acid sequences with structural features is developed and provides a modular prototype for follow-up, more extensive protein modeling, including larger proteins and sequence of variable sizes.

Carl David Jasper Causin, M. Fyta · 0 citations
Open access Sep 2026

Supervised Protein Structure Classification Using Topological Persistence With DeltaFold.

Recent advances in protein structure prediction have considerably increased the number of available structures, underscoring the need for scalable and accurate methods to compare and classify this huge amount of data. Current approaches based on sequence alignment or structural superposition often lack sensitivity and/or are computationally expensive. Here, we introduce the DeltaFold Classifier (DFC), a fast, alignment-free, protein structure classification pipeline based on topological data analysis. Protein three-dimensional structures are encoded using fixed-length vectors derived from persistent homology applied to point clouds of their spatial representation. These vectors, known as Biotopological Markers (BTMs), capture the intrinsic topological features of protein structure geometry-such as cycles and cavities-across multiple spatial scales. BTMs are invariant under translations and rotations, thereby avoiding the need for structural superposition. The DFC pipeline uses the features of these BTMs to train machine learning models for protein structure classification at various hierarchical levels defined in the SCOP and CATH classifications. It achieves performance comparable to that of structure-based comparison methods while substantially improving computational efficiency. It also outperforms sequence-based methods in tasks involving distant homology detection. In-depth analyses confirm that most BTM features contribute significantly to classification performance. These findings highlight the potential of topological vectorisations in structural bioinformatics and support the DFC pipeline as an effective and scalable tool for automated protein structure classification and annotation.

Unknown authors · 0 citations
Open access Jul 2026

MPLID (Membrane Protein–Lipid Interaction Database): A Large-Scale Experimental Resource of Residue-Level Protein–Lipid Contacts

Abstract Membrane proteins constitute approximately 20–30% of all proteomes and represent over 60% of current drug targets. Although protein–lipid interactions play important structural and regulatory roles in membrane-associated proteins, most existing structural resources focus on identifying whether a residue lies within a membrane region, typically inferred from computational hydrophobicity-based positioning algorithms. This approach does not directly address a distinct biological question: which residues at the protein surface make direct physical contact with lipid molecules? Answering this question from experimental data is critical for understanding lipid-mediated allostery, designing lipid-mimetic therapeutics, and training accurate machine learning models for lipid-binding-site prediction. We present MPLID (Membrane Protein–Lipid Interaction Database), a curated residue-level dataset comprising 4,704 membrane proteins representing 813 sequence clusters at 30% identity, 8,055,325 residues, and 80,439 structurally observed lipid-contact annotations (1.00% observed positive rate). Labels are derived exclusively from experimentally resolved lipid molecules in experimentally determined Protein Data Bank structures using a 4.0 Å all-atom heavy-atom distance cutoff. Because many native lipid interactions are lost or remain unresolved during purification and structure determination, this observed rate represents a lower bound, and the non-contact class inevitably contains false negatives. The dataset uses a curated list of 117 candidate lipid identifiers across ten functional categories, including 90 PDB-derived ligand codes audited against the RCSB Chemical Component Dictionary and 27 CHARMM-style lipid identifiers encountered in cryo-EM depositions. These identifiers span phospholipids, cardiolipin, sphingolipids, sterols, fatty acids, glycerolipids, detergent mimetics, lipid A components, and CHARMM simulation nomenclature. To minimize data leakage, proteins are clustered at 30% sequence identity using MMseqs2, yielding 813 clusters partitioned into training (2,578), validation (1,051), and test (1,075) splits. Amino acid composition analysis reveals biologically consistent enrichment at lipid-contact sites: tryptophan (1.88×), arginine (1.44×), glycine (1.36×), lysine (1.33×), and phenylalanine (1.23×) are enriched, whereas proline (0.51×), isoleucine (0.57×), and aspartate (0.59×) are depleted. MPLID addresses a distinct biological question from existing resources such as OPM, MemBlob, and BioDolphin/PLIP by identifying residues that directly contact experimentally resolved lipid molecules rather than residues positioned within computationally defined membrane boundaries. With 4,704 proteins and more than 8 million annotated residues, MPLID provides the scale needed for training deep-learning models for lipid-contact prediction, with applications in structure-guided drug design and membrane-protein engineering. The dataset adheres to FAIR principles and is freely available under a CC0 public-domain dedication. Structurally resolved contacts represent only a subset of biological protein–lipid interactions, and MPLID is intended as an experimentally grounded resource rather than a complete catalogue of lipid-binding sites.

F. B. Omage, Goran Neshich · 0 citations
Open access Aug 2026

SpY-C: Supervised Learning of Phosphopeptide Sequence Constraints Enables Global Prediction of SH2 Domain Binding

Tyrosine kinase signaling for cell development and homeostasis in multicelluar organisms and a major biochemical contribution is by driving interactions between phosphorylated tyrosines (pY) and SH2 domain containing proteins. This assembly is so important to driving cell outcomes that a wide variety of experimental and computational approaches have been used to understand which SH2-pY interactions occur, which still remains a challenge given the immensity (more than 45,000 pY and 120 SH2 domains in the human proteome). Based on biophysical constraints suggested by comprehensive contact mapping, here, we ask whether an approach might consider first asking if pY sequences conform to the shared rules of SH2 domain recognition by developing a classification approach that combines diverse training data. A wide range of validation suggests this approach, SpY-C, can classify pY sites as having the potential, or not, to be involved in SH2 domain interactions. We find that a relatively small set of representative SH2 binders, integrated from different experimental techniques, provides good classification. We use this classifier to annotate the human phosphoproteome and individual experiments, to explore the consequences of using super-SH2 domain reagents for pY enrichment, and to analyze the effects of mutations in altering pY site function. SpY-C provides a helpful step to more rapidly annotating pY function and for possibly improving machine learning approaches focused on specific SH2-pY interactions downstream of a first pass classification approach.

Alekhya Kandoor, A. C. Silva Oliveira, K. Machida et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.