Skip to content
Open access

WaterFlow: Prediction of Ordered Water Molecule Positions on Protein Structures

Aug 2026 · bioRxiv · 0 citations
Biology

TL;DR

WaterFlow is introduced, a flow-matching-based generator model and confidence model for predicting the positions of ordered water molecules in protein structures that outperforms the existing state of the art at every precision level and quantifies the tradeoff between data quantity and data quality.

Abstract

Ordered water molecules mediate many protein functions including stability, ligand binding, and catalysis. Predicting their positions with sub-angstrom accuracy would support protein design, binding affinity prediction, and automated model building in X-ray crystallography and cryo-EM. However, water molecule prediction lags behind protein and other molecule structure predictions. We introduce WaterFlow, a flow-matching-based generator model and confidence model for predicting the positions of ordered water molecules in protein structures. WaterFlow outperforms the existing state of the art at every precision level. We demonstrate that this model not only predicts ground truth modeled water molecules, including those around ligands, but also fits the underlying experimental data well, and therefore proposes that it may be used for both prediction and modeling water molecules. We demonstrate that WaterFlow’s novel predictions are often associated with positive electron difference density, meaning the model places water molecules at sites the original structure depositions omitted. We use this improved model to address the data constraint. By mapping the Pareto front of achievable accuracy of water molecule prediction, alongside analysis of different training data schemas, we quantified the tradeoff between data quantity and data quality, demonstrating the diversity of high quality structures is limiting the results possible. Overall, WaterFlow predicts ordered water to serve as a solvent module for structure-based drug design, and predicted structures, as well as for water molecule placement during crystallographic refinement.

Read PDF

Similar papers

Does Surface Conservation Yield? Application to Data-Driven Docking

The interface prediction program WHISCY is presented, which combines surface conservation and structural information to predict protein–protein interfaces and demonstrates the potential of using interface predictions to drive protein–protein docking.

S. D. de Vries, A. V. van Dijk, A. M. J. J. Bonvin · 0 citations
Open access Aug 2026

PreFold-dG: estimating binding affinity of protein–protein interaction from intermediate representations of protein folding model

PreFold-dG is presented, a model that estimates binding affinities of protein complexes utilizing intermediate embeddings from Boltz-2, an open-source foundation model for protein structure prediction and achieved state-of-the-art performance on well-established binding affinity prediction benchmarks and demonstrated robustness on independent test sets.

Sungjoon Park, Soorin Yim, Dongyun Kim et al. · 0 citations
Open access Aug 2026

Conserved water molecules shape the pathogenicity of missense variants in human proteins.

Conserved water molecules (CWMs) are tightly bound solvent molecules that occupy well-defined, recurrent positions in protein structures. Although they are known to influence protein stability, function, and ligand binding, their role in shaping the effects of human missense variants remains largely unexplored. Here, we demonstrate that CWMs are a previously underappreciated determinant of missense variant pathogenicity. By predicting ligand-binding and CWM sites across human PDB structures and mapping missense variants to these sites and the remaining protein surface, we found that pathogenic variants were significantly enriched at CWM sites, whether overlapping or outside other ligand-binding regions. This enrichment exceeded that observed for binding sites as a whole, indicating a broader role for water-mediated interactions in modulating variant effects. To explore a mechanistic basis for this association, we performed molecular dynamics simulations of human lysosomal acid glucosylceramidase (GCase), encoded by GBA1 and implicated in Gaucher disease and Parkinson's disease risk. Selective destabilization of a CWM site in wild-type GCase produced structural and dynamical changes resembling those observed in the pathogenic L444P variant, whereas stabilization of this site in L444P shifted several measures toward wild-type behavior. These results suggest that disruption of a single CWM can contribute to long-range structural remodeling observed in a disease-associated variant. Together, our findings identify CWMs as a novel structural constraint shaping the distribution and effects of pathogenic missense variants. Incorporating water-mediated interactions into structural models provides a generalizable framework for interpreting human genetic variation and its contribution to disease.

Janez Konc, Karmen Recer, Tanja Kunej et al. · 0 citations
#protein folding Open access Aug 2026

Predictive all-atom simulations of disordered proteins and biomolecular condensates through osmometry-guided force-field optimization

All-atom simulations with explicit solvent provide the most detailed and accurate description of dynamics and mechanisms in intrinsically disordered proteins and their condensates. However, interactions involving charged residues and ions remain a persistent source of systematic error. Here we introduce an osmometry-guided optimization strategy that directly targets residue–residue, residue–ion and ion–ion interactions. Osmotic pressure provides key experimental information on molecular interactions and can be calculated directly and rapidly from simulations, enabling efficient iterative force-field optimization. The resulting parameters improve agreement of all-atom simulations with a range of experimental data: single-molecule FRET measurements for 16 monomeric intrinsically disordered regions; NMR relaxation data for a complex between an IDP and a folded protein domain; and mean FRET efficiencies and chain reconfiguration times of IDPs in biomolecular condensates of highly charged proteins. For such condensates, simulations with an osmometry-calibrated force field provide the missing link for predicting condensate dynamics across length and time scales. The presented optimization strategy is broadly extensible to other interaction classes, including those governing protein–DNA and protein–RNA assemblies.

Miloš T. Ivanović, Valentin von Roten, Benjamin Schuler et al. · 0 citations
Review Open access Jul 2026

Solvent Interaction Analysis: A New Lens for Protein Structure and Diagnostics

The data support aqueous solvent interaction analysis as a broadly applicable, mechanistically grounded technology for protein characterization, drug–protein interaction studies, and structure-centric biomarker development, exemplified by the clinical translation of IsoPSA.

B. Zaslavsky, M. Stovsky, V. Uversky · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.