Skip to content
Open access

MEM-CALVADOS: A Residue-Level Model for Flexible Proteins at Membrane Interfaces

Aug 2026 · bioRxiv · 0 citations · 1 references
Biology

TL;DR

The model provides a computationally efficient framework for studying conformational ensembles and assembly of proteins at bilayer–water interfaces and is validated against Wimley–White free energies of transfer of hydrophobic peptides and against an experimentally refined conformational ensemble of a flexible membrane receptor.

Abstract

Many membrane proteins contain intrinsically disordered regions (IDRs) that play key biological roles by providing structural plasticity, harboring sites for post-translational modifications, and mediating protein clustering and phase separation. Residue-level molecular models parameterized against experimental data have provided insights into how IDR sequence controls conformational properties and phase behavior in soluble proteins. Here, we extend this modeling framework to membrane-associated IDRs. We adapt a coarse-grained lipid model, iSoLFv2, for four phospholipids and combine it with CALVADOS, a residue-level model for IDRs and multi-domain proteins. Protein–lipid cross-interaction parameters are calibrated to reproduce predicted insertion and orientation of transmembrane proteins with diverse architectures. We validate the resulting model against Wimley–White free energies of transfer of hydrophobic peptides and against an experimentally refined conformational ensemble of a flexible membrane receptor. Finally, we show the applicability of the model to a membrane-associated assembly of signaling proteins. The model provides a computationally efficient framework for studying conformational ensembles and assembly of proteins at bilayer–water interfaces.

Read PDF

Similar papers

Open access Mar 2026

A membrane insertion code for intrinsically disordered proteins

Membrane association of intrinsically disordered proteins (IDPs) mediates various cellular functions including membrane remodeling and signal transduction. Whereas membrane association through amphipathic helices and polybasic motifs is well understood, sequence determinants for the insertion of aromatic residues into the membrane hydrophobic core are still poorly characterized. Here, we decipher the sequence code for membrane insertion of aromatic-centered motifs. For an initial set of 10 9-residue aromatic-centered sequences, all-atom molecular dynamics simulations and the positioning of proteins in membranes (PPM) method produced very similar membrane insertion propensities. Applying PPM to a full library of 1.2 × 106 sequences with an F, W, or Y residue flanked by L, R, G, N, or E at four positions on either side, we found that aliphatic (L) and basic (R) residues favor membrane insertion, whereas acidic (E) and polar (N) residues disfavor it. Guided by these rules, we developed a mathematical model dubbed AroMIP (Aromatic Membrane Insertion Predictor) to predict the membrane insertion propensities of aromatic-centered motifs. AroMIP achieves 91.2%, 92.0%, and 99.7% accuracies for F-, W-, and Y-centered motifs, respectively, in disordered regions of the human proteome and is available as a web server at https://zhougroup-uic.github.io/AroMIP/. The present work provides the sequence basis and a mechanistic understanding of how IDPs employ aromatic-centered motifs to drive membrane insertion, and enriches the tools for the study of IDP-membrane association.

Fidha Nazreen Kunnath Muhammedkutty, Huan-Xiang Zhou · 1 citation
Sep 2026

A Biophysical Hypothesis for Compartment-Specific Conformational States of the IQSEC2 Intrinsically Disordered Protein.

The postsynaptic density (PSD) of neuronal synapses is a crowded, viscous membraneless compartment consisting of densely packed protein mixtures formed and maintained through liquid-liquid phase separation. A key regulator of synaptic function within the PSD is IQSEC2, an intrinsically disordered protein (IDP), which functions as a guanine nucleotide exchange factor (GEF), promoting the activation of the small GTPase ARF6 by catalyzing the exchange of ARF6 bound GDP for GTP. Experiments have shown that IQSEC2 is inactive in a folded state in the dendritic cytosol but can transiently adopt a catalytically active extended conformation in the PSD upon activation by neurotransmitter mediated calcium influx. In this study, we performed accelerated molecular dynamics (aMD) simulations of the IQSEC2 folding process in aqueous solvent and applied the Gibbs ergodic hypothesis formulated in statistical ensembles for post-processing and analysis of the results. We compared the effects of two sets of parameters on folding: a force field and a water model. The OPC water model optimized for proteins with disordered structure in the extended state in bulk aqueous solvent predicted an incorrect distorted 3D folded structure of IQSEC2. We propose that, due to fundamental biophysical differences between the dendritic cytosol and the postsynaptic density, there are two distinct classes of IDPs functioning in these different environments. These differences should be considered when optimizing force fields and parameters of water models for studying protein folding.

M. Shokhen, A. Albeck, N. Levy et al. · 0 citations
Open access Aug 2026

Lipid-Shell PARCH: A Physically Motivated Scale for Transmembrane Residue Hydropathy

Understanding how amino acid residues partition from water into lipid bilayers is fundamental to membrane protein folding, stability, and function. Existing hydrophobicity scales derive from isolated protein systems, organic solvent approximations, or computationally expensive free energy methods, each with significant limitations in capturing the full thermodynamic and topographic complexity of membrane protein environments. Here we present a lipid-shell modification to the protocol for assigning a residue’s character on a hydropathy (PARCH) scale, in which the protein’s first hydration shell is enclosed by a lipid boundary layer during thermal annealing. This modification preserves the core PARCH methodologyevaluating water retention around residues as a function of temperaturewhile imposing the chemical potential boundary condition appropriate for membrane-embedded proteins. Using the OmpLA host–guest system, we compute PARCH values (PVs) that align with the experimental water-to-bilayer transfer free energy scale of Moon and Fleming (PNAS, 108, 10174–10177, 2011) without calibration to that data. We further demonstrate that depth-dependent PV profiles for arginine and leucine mirror experimentally measured partition energies across six membrane positions. Finally, PVs for a tandem arginine double mutant reveal a per-residue cooperative hydration redistribution that provides microscopic insight into thermodynamic cooperativity previously measured experimentally. Together, these results establish the lipid-shell PARCH modification as a computationally affordable and physically meaningful approach to quantifying membrane hydropathy that is sensitive to residue identity, membrane depth, and cooperative hydration among tandem charged residues.

Ratnakshi Mandal, S. Nangia · 0 citations
Open access Aug 2026

Benchmarking AI-generated structural ensembles of membrane proteins against physics-based modelling

It is demonstrated that BioEmu can generate plausible conformational ensembles for relatively large, six-and seven-pass membrane proteins, sampling rare states at a fraction of the computational cost of conventional MD simulations, suggesting that AI-based ensemble generation could provide an accessible approach for exploring membrane protein dynamics and complement conventional molecular modelling approaches.

B. Clifton, Adam G Grieve, Robin A. Corey · 0 citations
Open access Jul 2026

The topology of transmembrane protein-protein interaction interfaces is encoded in their physicochemical features

Transmembrane (TM) protein-protein interactions (PPIs) are essential mediators of signal transduction, transport of solutes and communication, yet the biophysical features that characterize their diverse topologies remain poorly understood. To reduce this knowledge gap we use the human solute carrier (SLC) interactome to study whether physicochemical features of PPI interfaces encode information about their TM topology (i.e., their position and arrangement with respect to the membrane). To this end we predicted structures for 2, 055 experimentally validated PPIs using AlphaFold v3.0, performed molecular dynamics simulations and annotated the PPI interfaces by TM coverage. As a result every interface is characterized by 63 physicochemical, structural and energy features. A reproducible machine learning workflow allows us to study the interdependence between these interface properties and interface topology. We found that amino acid composition and secondary structure contributed most to distinguishing soluble from fully membrane-embedded interfaces. Membrane-embedded interfaces contained a larger fraction of hydrophobic residues and α-helices, while charged residues were depleted. For partially and full membrane-embedded interfaces, amino acid composition and secondary structure get more similar and differences are increasingly observed in charge and energy related features. As particularly noteworthy we find that PPI interface characteristics vary with the number of annotated TM segments mainly through differences in proportions of secondary structure, charge, and flexibility. We hence conclude that PPI interface characteristics harbor substantial information about TM interface topology and provide a framework for the study and design of membrane protein interaction interfaces.

Lisa Allmesberger-Riegler, Fabian Frommelt, Brianda L. Santini et al. · 0 citations
Open access Jul 2026

MPLID (Membrane Protein–Lipid Interaction Database): A Large-Scale Experimental Resource of Residue-Level Protein–Lipid Contacts

Abstract Membrane proteins constitute approximately 20–30% of all proteomes and represent over 60% of current drug targets. Although protein–lipid interactions play important structural and regulatory roles in membrane-associated proteins, most existing structural resources focus on identifying whether a residue lies within a membrane region, typically inferred from computational hydrophobicity-based positioning algorithms. This approach does not directly address a distinct biological question: which residues at the protein surface make direct physical contact with lipid molecules? Answering this question from experimental data is critical for understanding lipid-mediated allostery, designing lipid-mimetic therapeutics, and training accurate machine learning models for lipid-binding-site prediction. We present MPLID (Membrane Protein–Lipid Interaction Database), a curated residue-level dataset comprising 4,704 membrane proteins representing 813 sequence clusters at 30% identity, 8,055,325 residues, and 80,439 structurally observed lipid-contact annotations (1.00% observed positive rate). Labels are derived exclusively from experimentally resolved lipid molecules in experimentally determined Protein Data Bank structures using a 4.0 Å all-atom heavy-atom distance cutoff. Because many native lipid interactions are lost or remain unresolved during purification and structure determination, this observed rate represents a lower bound, and the non-contact class inevitably contains false negatives. The dataset uses a curated list of 117 candidate lipid identifiers across ten functional categories, including 90 PDB-derived ligand codes audited against the RCSB Chemical Component Dictionary and 27 CHARMM-style lipid identifiers encountered in cryo-EM depositions. These identifiers span phospholipids, cardiolipin, sphingolipids, sterols, fatty acids, glycerolipids, detergent mimetics, lipid A components, and CHARMM simulation nomenclature. To minimize data leakage, proteins are clustered at 30% sequence identity using MMseqs2, yielding 813 clusters partitioned into training (2,578), validation (1,051), and test (1,075) splits. Amino acid composition analysis reveals biologically consistent enrichment at lipid-contact sites: tryptophan (1.88×), arginine (1.44×), glycine (1.36×), lysine (1.33×), and phenylalanine (1.23×) are enriched, whereas proline (0.51×), isoleucine (0.57×), and aspartate (0.59×) are depleted. MPLID addresses a distinct biological question from existing resources such as OPM, MemBlob, and BioDolphin/PLIP by identifying residues that directly contact experimentally resolved lipid molecules rather than residues positioned within computationally defined membrane boundaries. With 4,704 proteins and more than 8 million annotated residues, MPLID provides the scale needed for training deep-learning models for lipid-contact prediction, with applications in structure-guided drug design and membrane-protein engineering. The dataset adheres to FAIR principles and is freely available under a CC0 public-domain dedication. Structurally resolved contacts represent only a subset of biological protein–lipid interactions, and MPLID is intended as an experimentally grounded resource rather than a complete catalogue of lipid-binding sites.

F. B. Omage, Goran Neshich · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.