Skip to content

Quantifying the Peripheral Surface Information Entropy from Conformational Ensembles of Globular Protein-Peptide Complexes.

Jul 2026 · Biophysical Journal · 0 citations
Medicine

Abstract

Predicting favorable protein-peptide binding events remains a central challenge in biophysics, with continued uncertainty surrounding how nonlocal effects shape the global energy landscape. Here, we introduce peripheral surface information (PSI) entropy, SΨ, a quantitative measure of the statistical variability in apolar and charged non-interacting surface (NIS) proportions across conformational ensembles. Within the Gibbs free-energy relation ΔG = ΔH - TΔS, SΨ is proposed as a computationally tractable entropic proxy rather than a direct thermodynamic observable or stand-alone estimator of binding affinity. Using energy-directed molecular docking via HADDOCK3 and explicit-solvent molecular dynamics simulations, it is demonstrated that favorable binding partners exhibit emergent, low-entropy N-states (discrete macrostates in NIS state space) indicative of preferential apolar/charged surface configurations. Across dozens of peptides and multiple receptor systems (WW, PDZ, and MDM2 domains), dominant N-states persisted under varied docking parameters and initial conditions. A meta-ensemble of 657 complexes from 36 experiments over 15 years confirmed the presence of dominant NIS modes independent of in silico methodology, suggesting an evolutionary selection pressure toward specific NIS fingerprints. These findings establish SΨ as a thermoinformatic descriptor that encodes favorable binding constraints into unique statistical signatures of the NIS.

View source

Similar papers

Preprint Aug 2026

PHASE: encoding global protein ensembles with local Hamiltonians and all-atom backmapping

Protein function is governed by conformational ensembles, which can be viewed as high-dimensional probability distributions over molecular conformations. Yet the statistical organization of these distributions is often represented only implicitly, either through collections of simulation trajectories or within high-capacity generative models. Here, we introduce PHASE (Protein Hamiltonians for Sampling of Ensembles), a system-specific framework that converts atomistic conformational ensembles into an explicit and interpretable statistical model. Applied to ten conformational ensembles derived from approximately 37$\mu$s of atomistic simulations of the adenosine A2A receptor, Hamiltonians containing only local residue couplings within 6$\mathring{A}$ reproduce residue-wise and pairwise microstate statistics, including correlations between residues that are not directly coupled in the model. Moreover, independently fitted inactive and active reference Hamiltonians define an endpoint preference coordinate that organizes newly sampled ligand-, effector- and conformation-dependent ensembles along the A2A activation landscape without receiving these biochemical labels as model inputs. Finally, a cluster-conditioned all-atom reconstruction model preserves the prescribed residue microstate patterns of newly sampled configurations, closing the coarse-graining-sampling-backmapping cycle. The resulting discrete representation additionally admits direct QUBO encoding, enabling classical annealing and providing a route toward future quantum-annealing implementations. PHASE therefore provides a protein-general procedure for constructing compact, interpretable and atomistically realizable statistical models of protein conformational ensembles.

Daniele Angioletti, Marco S. Nobile, Matteo Carli et al. · 0 citations
Open access Aug 2026

Pi-Ensemble: Sequence-guided generation of interpolated protein conformational ensembles

Proteins are critical biomolecular machines that populate ensembles of interconverting conformations. Many biological processes depend on transitions between metastable states. Although molecular dynamics (MD) simulations provide a physically grounded route to characterize these motions, routine sampling of large-scale conformational transitions remains computationally demanding. Recent advances in protein structure prediction have created new opportunities for ensemble generation, but many existing approaches require noising inputs, task-specific training, supervised fitting on extensive MD data, or experimentally-informed restraints. Here, we introduce Pi-Ensemble (Predicting Interpolated Ensemble), a sequence-guided framework for generating protein conformational ensembles interpolating between two structural anchor states. Unlike previous methods, Pi-Ensemble alternately leverages inverse-folding and structure-prediction models to propose intermediate conformations between known protein states, generating diverse ensembles without additional training. We evaluate Pi-Ensemble across diverse protein systems, including enzymes, transporters, receptors, and benchmark cases with reference MD simulations or experimental Double Electron-Electron Resonance (DEER) data. Pi-Ensemble recovers physically plausible intermediate conformations, captures transition pathways observed in large-scale MD simulations, and generates structures consistent with experimental distance distributions. Furthermore, Pi-Ensemble-generated conformations provide effective starting seeds for parallel MD simulations, improving conformational exploration and accelerating convergence relative to simulations initiated only from endpoint structures. These results establish sequence-guided structural interpolation as a practical strategy for probing protein conformational landscapes. By generating diverse and physically reasonable conformational proposals without long-timescale MD or model retraining, Pi-Ensemble provides an extensible framework for studying protein flexibility, guiding adaptive sampling, and accelerating mechanistic investigations of protein function.

Hassan Nadeem, D. Kleiman, Yuming Zhou et al. · 0 citations
Preprint Aug 2026

A Quantum Circuit Framework for Protein Ensemble-Level Energetics

Proteins occupy heterogeneous free-energy landscapes in which high-entropy ensembles converge toward compact, low-energy basins with multiple sub-states. Molecular dynamics can access these landscapes at atomic resolution, but exhaustive sampling remains computationally demanding. Meanwhile, most quantum approaches target only single optimal structures, leaving full ensemble energetic heterogeneity unexplored. We introduce a residue-level, gate-based quantum circuit framework for coarse-graining protein thermodynamics. Each amino acid is represented as a two-state qubit (stabilised vs. excited solvation state) based on residue solvation energetics. A structure-informed entanglement block then encodes covalent and non-covalent contacts using parameterised controlled gates, embedding correlations across the residue-interaction network. Sampling the circuit ($\sim 10^6$ measurements) yields binary thermodynamic microstates used to compute protein energy distributions, residue-level statistical couplings, energetic sensitivities, and information gains relative to total free energy. We showcase the framework on the benchmark Trp-cage miniprotein 1L2Y (TC5b) and 9GDL, a disulfide-stabilised Trp-cage-fortified exenatide chimera. For 1L2Y, the circuit reproduces a structured, folding-funnel-like energy distribution. Comparative analysis with 9GDL reveals shifts in global energy distributions and residue-level stability profiles. Coupling and information-theoretic analyses localise residues associated with ensemble reorganisation, while multi-body couplings show the circuit resolves both direct and indirect statistical correlations. This framework expands quantum protein modelling beyond single-structure optimisation toward ensemble-level characterisation, capturing key features of rugged energy landscapes to guide protein design, mutation mapping, and allosteric pathway identification.

Pratik Patil, Bhushan Bonde, B. Choubey · 0 citations
Open access Jul 2026

Integrative Ensemble Modeling reveals RNA conformations targetable by small molecules

RNA molecules explore heterogeneous conformational ensembles that are essential for their biological function and molecular recognition, yet this intrinsic flexibility poses a major challenge for structure-based drug discovery. In particular, the absence of well-defined binding pockets in static structures limits the identification of ligandable sites. Here, we present an integrative ensemble-based approach that combines enhanced-sampling molecular dynamics simulations with Nuclear Magnetic Resonance data to characterize the conformational landscape of the HIV-1 TAR RNA at atomic resolution. Starting from extensive sampling, we refined the resulting conformational distribution through maximum-entropy reweighting to achieve quantitative agreement with experimental data. Analysis of the reweighted ensemble reveals a diverse set of conformational substates, including compact arrangements that exhibit pocket features compatible with ligand recognition and overlap with known ligand-bound structures. At the same time, highly ligandable conformations, which are only marginally populated, might nonetheless be critical for RNA recognition. Our results demonstrate that integrative ensemble modeling can reveal pharmacologically relevant RNA conformations that are not apparent from experimental static structures, providing a framework for ensemble-based strategies in RNA-targeted drug discovery.

Stefano Bosio, Vincent Schnapka, Mattia Bernetti et al. · 0 citations
Open access Aug 2026

Benchmarking AI-generated structural ensembles of membrane proteins against physics-based modelling

Proteins dynamically switch between a continuum of interconverting conformational states, and understanding these structural dynamics is important for understanding protein function and for developing therapeutics. Molecular dynamics (MD) simulations can provide insight into protein conformational ensembles, but sampling rare conformational states can require substantial computational resources. The recent development of AI-based approaches for generating protein conformational ensembles, such as the Biomolecular Emulator (BioEmu), offers a potential alternative, although it remains unclear whether these approaches can accurately capture the conformational landscapes, especially for special cases such as membrane proteins. Here, we assess the ability of BioEmu to model the conformational dynamics of a model membrane protein, the bacterial rhomboid intramembrane proteases GlpG. We find that BioEmu generates a range of conformations corresponding to both open and closed states of the rhomboid lateral gate, including states associated with different stages of the catalytic cycle. These conformations broadly correspond to states sampled during microsecond-timescale MD simulations, although BioEmu does not reproduce the full conformational landscape observed using MD. BioEmu also samples substantial conformational heterogeneity within the soluble domains of rhomboids, which are highly flexible and poorly represented in experimental structures. Overall, our findings demonstrate that BioEmu can generate plausible conformational ensembles for relatively large, six-and seven-pass membrane proteins, sampling rare states at a fraction of the computational cost of conventional MD simulations. These results suggest that AI-based ensemble generation could provide an accessible approach for exploring membrane protein dynamics and complement conventional molecular modelling approaches.

B. Clifton, Adam G Grieve, Robin A. Corey · 0 citations
Open access Jul 2026

Observation of Mechanical and Kinetic Distinctions between Individual Isoleucine and Arginine Residues in a Peptide Dissociating from a Model Lipid Bilayer

Peptide–lipid membrane interactions underlie many essential biological processes, yet the molecular determinants of peptide partitioning and dissociation from lipid bilayers remain incompletely understood. Here, we combine coarse-grained molecular dynamics (CG MD) simulations and atomic force microscopy (AFM)-based force spectroscopy to study the structural dynamics, energetics, and kinetics of penta-X5 peptides (WLLLX, with X = R or I) interacting with POPC bilayers. To elucidate how the position and identity of a single guest residue X modulates peptide–membrane interactions, we present these results in the context of the canonical Wimley–White penta-X (WLXLL) motif. Our findings from CG simulations are consistent with penta-X5 peptides adopting snorkeling conformations beneath the bilayer surface and with a dissociation scenario in which the final two or three residues detach almost simultaneously under the applied pulling force. Ensemble analyses of the reconstructed potential of mean force profiles lead to multiple energetic dissociation pathways. In combination with kinetic modeling of AFM rupture force distributions, the data reveal that both the mechanics (dissociation force) and kinetics (off rate) of peptide detachment are sensitive to the identity and sequence position of individual residues. These results highlight the power of integrating CG MD and single-molecule force spectroscopy to unravel residue-specific, sequence-dependent factors underlying peptide–lipid interactions.

Ryan S. Smith, Krishna P Sigdel, D. R. Weaver et al. · 0 citations