Skip to content
Open access

Variant characterization in the intrinsically disordered human proteome

Jul 2026 · Nature Structural & Molecular Biology · Vol 33, pp. 1183 - 1193 · 0 citations · 84 references
Medicine

TL;DR

Proteome-wide prediction and structural modeling of disordered protein interaction interfaces advance characterization of disease-associated variants in disordered protein regions.

Abstract

Variant effect prediction remains a key challenge in precision medicine. Computational models are increasingly successful in the characterization of missense variants in folded protein regions. However, 37% of all annotated missense variants reside in the 25% of the proteome that is intrinsically disordered, lacking positional sequence conservation and stable structures. To advance the characterization of variants in intrinsically disordered protein regions (IDRs), we combined sequence pattern searches with AlphaFold to structurally annotate 1,300 protein–protein interactions with interfaces mediated by short disordered motifs binding to folded domains in partner proteins. These interfaces were selected based on their overlap with uncertain missense variants enabling structural model-based prediction of deleterious effects of 1,187 of these variants in IDRs. Extensive experimental efforts validated the predicted interfaces and deleterious variant effects that were predicted as benign by AlphaMissense, demonstrating that the combination of sequence analysis and structural modeling can readily generate numerous testable hypotheses of variant effects on protein function in IDRs. Proteome-wide prediction and structural modeling of disordered protein interaction interfaces advance characterization of disease-associated variants in disordered protein regions.

Read PDF

Similar papers

Open access Aug 2026

A discrete protein subset drives structure prediction discordance in orphan proteins

The discordance persists: pLDDT correlates positively with PUNCH2 disorder in random and de novo proteins and negatively with β-strand fraction, opposite to the conserved and disordered baselines, a concrete failure mode that protein designers and other working on sequences remote in sequence space should be aware of when relying on predictor outputs.

Lars A. Eicholt, Lasse Middendorf · 0 citations
Open access Aug 2026

Conserved water molecules shape the pathogenicity of missense variants in human proteins.

Conserved water molecules (CWMs) are tightly bound solvent molecules that occupy well-defined, recurrent positions in protein structures. Although they are known to influence protein stability, function, and ligand binding, their role in shaping the effects of human missense variants remains largely unexplored. Here, we demonstrate that CWMs are a previously underappreciated determinant of missense variant pathogenicity. By predicting ligand-binding and CWM sites across human PDB structures and mapping missense variants to these sites and the remaining protein surface, we found that pathogenic variants were significantly enriched at CWM sites, whether overlapping or outside other ligand-binding regions. This enrichment exceeded that observed for binding sites as a whole, indicating a broader role for water-mediated interactions in modulating variant effects. To explore a mechanistic basis for this association, we performed molecular dynamics simulations of human lysosomal acid glucosylceramidase (GCase), encoded by GBA1 and implicated in Gaucher disease and Parkinson's disease risk. Selective destabilization of a CWM site in wild-type GCase produced structural and dynamical changes resembling those observed in the pathogenic L444P variant, whereas stabilization of this site in L444P shifted several measures toward wild-type behavior. These results suggest that disruption of a single CWM can contribute to long-range structural remodeling observed in a disease-associated variant. Together, our findings identify CWMs as a novel structural constraint shaping the distribution and effects of pathogenic missense variants. Incorporating water-mediated interactions into structural models provides a generalizable framework for interpreting human genetic variation and its contribution to disease.

Janez Konc, Karmen Recer, Tanja Kunej et al. · 0 citations
Open access Aug 2026

Structure-Based Network Analysis of AlphaFold Structure Predictions Identifies Putative Causative Variants of Inherited Retinal Disease

SBNA can identify variants in human proteins that are likely to cause disease, and it can help predict variants causative of IRDs in an unbiased fashion using both AlphaFold2-generated structural models and experimental structural data.

Blake Hauser, E. Place, Yuyang Luo et al. · 0 citations
Open access Jul 2026

Interpretable Prediction of Phase Separation and Disease Variant Effects in Intrinsically Disordered Regions

An interpretable ensemble machine-learning framework that integrates protein language model embeddings of sequence and predicted structure to predict LLPS propensity and classify proteins as self-separating or partner-dependent and identifies critical phase-separating regions and quantifies mutation-induced perturbations in LLPS.

Mingjie Zhao, Sushant Kumar · 0 citations
Open access Jul 2026

Computational protein stability analysis of SCN1A missense variants reveals domain‐dependent stability patterns

Abstract Objective To determine whether computational protein‐stability predictions discriminate pathogenic from benign SCN1A missense variants, and to characterize the structural distribution of predicted destabilization among pathogenic variants. Methods On an AlphaFold3‐predicted Nav1.1 structure, FoldX, and Rosetta Cartesian ΔΔG were computed for a single ClinVar snapshot of pathogenic/likely‐pathogenic (P/LP) and benign/likely‐benign (B/LB) missense variants and its extension to ClinVar variants of uncertain significance (VUS) and gnomAD v4.1 variants; membrane‐aware RosettaMP was applied to the patch‐clamp subgroup. Pathogenic variants were stratified by functional domain. Results Pathogenic variants were more destabilizing than benign (FoldX 2.61 vs. 0.31 kcal/mol, p = 1.27 × 10−11; ROC‐AUC = 0.760), concordant with Rosetta (ROC‐AUC = 0.697; ρ = 0.660). Destabilization was domain‐dependent: pore (P‐loop/selectivity‐filter) pathogenic variants were depleted of stability‐neutral variants (0.40‐fold; Bonferroni‐adjusted p = 4.3 × 10−7), whereas S4 voltage‐sensor variants were enriched for them (2.32‐fold; p = 0.013). Across ~3300 non‐redundant variants, gnomAD‐common variants resembled benign controls and VUS were intermediate (mean ΔΔG 0.98 kcal/mol; 18.5% strongly destabilizing), with the domain pattern preserved. Among 64 patch‐clamp variants, stability did not separate gain‐ from loss‐of‐function, though gain‐of‐function variants clustered in voltage‐sensing domains and were absent from the pore. Significance Computational stability analysis thus adds a mechanistic layer complementary to the conventional gating‐dysfunction view, distinguishing a destabilized pore‐region subset—for which proteostasis impairment is a candidate, though unproven, mechanism—from a structurally tolerated S4 subset whose pathogenicity is stability‐independent. As a hypothesis‐generating rather than mechanism‐defining approach, this stratification prioritizes candidate variants—including the 18.5% of VUS that are strongly destabilizing—for direct functional and surface‐expression validation in SCN1A‐related epilepsies. Plain Language Summary We used computational modeling to predict how thousands of SCN1A genetic variants influence the stability of the Nav1.1 sodium channel protein. Disease‐causing variants tended to destabilize the protein more than benign variants, and variants in the pore region—where ions flow through the channel—were predominantly destabilizing. This is consistent with loss‐of‐function arising from misfolding and degradation of the channel protein in this subset of variants. By contrast, variants in the voltage‐sensing region were often structurally tolerated, indicating that their disease‐causing effects likely arise through a different mechanism that requires direct functional measurement to define. Accordingly, the analysis nominates a candidate pore‐region subset potentially affected by proteostasis impairment and a complementary stability‐neutral subset warranting functional evaluation.

Y. Shim, E. Kang, Naeun Kwak et al. · 0 citations
Open access Jan 2026

Computational Characterization of Pathogenic LMNA Missense Variants: Structural Instability, Altered Binding, and Conformational Dynamics

Background Mutations in the LMNA gene underlie a broad spectrum of laminopathies, including muscular dystrophies, cardiomyopathies, and premature aging syndromes; however, the molecular mechanisms by which missense variants disrupt Lamin A structural integrity remain incompletely characterized. Systematic computational approaches for prioritizing pathogenic variants and elucidating their structural consequences are critically needed. Methods An integrated multistep in silico framework was employed to investigate the structural and functional consequences of LMNA missense variants. Variant prioritization was performed using the Evo2 nucleotide language model via delta log‐likelihood scoring, followed by bioinformatic annotation using SIFT, PANTHER‐PSEP, PhD‐SNP, and E‐SNPs&GO. Protein stability assessment was conducted with DynaMut, INPS‐MD, I‐Mutant2.0, and MUpro. Variants localized within globular domains—N456D, N456T, and G465D—together with the known pathogenic variant M540T as a positive control, were selected for three‐dimensional structural modeling using PyMOL and AlphaFold2, molecular docking with lonafarnib as a reference ligand via AutoDock Vina, and 100 ns molecular dynamics simulations using GROMACS with the Amber ff14SB force field. Conformational dynamics were characterized through principal component analysis and free‐energy surface construction. Results Evo2‐based screening of the full LMNA coding sequence identified 50 high‐priority loss‐of‐function variants, of which N456D, N456T, and G465D were retained for structural investigation based on their globular domain localization and multitool pathogenicity predictions. All three variants were consistently predicted to alter physicochemical properties and reduce structural stability relative to wild‐type Lamin A. Molecular docking revealed mutation‐dependent changes in lonafarnib binding profiles. The known pathogenic control M540T exhibited comparable structural and dynamic behavior, supporting the reliability of the prioritization workflow. Molecular dynamics analyses demonstrated altered RMSD trajectories, increased residue‐level flexibility, and modified hydrogen bonding patterns in mutant systems. Free‐energy landscape analyses revealed expanded conformational basins, particularly pronounced in the G465D variant, indicating increased structural plasticity. Conclusion This integrated computational framework provides a systematic strategy for prioritizing pathogenic LMNA variants and characterizing their structural consequences at the atomic level. The identified variants—N456D, N456T, and G465D—represent structurally disruptive substitutions consistent with the established role of globular domain destabilization in other laminopathy‐associated variants, offering testable hypotheses for experimental validation in cellular and animal models.

E. Aktaş, Ceren Nizamoğlu, Salvador Ventura · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.