Skip to content
Open access

A discrete protein subset drives structure prediction discordance in orphan proteins

Aug 2026 · bioRxiv · 0 citations · 103 references
Biology

TL;DR

The discordance persists: pLDDT correlates positively with PUNCH2 disorder in random and de novo proteins and negatively with β-strand fraction, opposite to the conserved and disordered baselines, a concrete failure mode that protein designers and other working on sequences remote in sequence space should be aware of when relying on predictor outputs.

Abstract

Structure and disorder predictors are increasingly used as decision-grade tools in protein engineering and in the analysis of newly emerged proteins, yet how the current state-of-the-art behaves on sequences outside the well-charted evolutionary space remains poorly characterised. We previously reported that AlphaFold2 confidence and the disorder predictor flDPnn produced discordant predictions for naturally evolved de novo Drosophila proteins and for shuffled sequences. Here, we revisit the comparison with AlphaFold3 and the best-performing disorder predictor PUNCH2 on the same sequence sets together with conserved Drosophila proteins and intrinsically disordered proteins. The discordance persists: pLDDT correlates positively with PUNCH2 disorder in random and de novo proteins and negatively with β-strand fraction, opposite to the conserved and disordered baselines. A class-specific, score-defined driver subset jointly captures the unusual high-pLDDT, high-disorder, low-strand combination and contains 24.5% of de novo, 29.4% of random, 5.1% of conserved, and 1.3% of disordered proteins. Removing this subset normalises the correlations. A held-out classifier trained on architectural and compositional features that were not used in the driver definition recovers the subset, with helix and coil fraction, sequence length, entropy and hydropathy as the strongest predictors. The discordance is therefore not a sequence-class artefact but a localised, compositionally identifiable phenotype that current predictors handle in a non-canonical way - a concrete failure mode that protein designers and other working on sequences remote in sequence space should be aware of when relying on predictor outputs.

Read PDF

Similar papers

Open access Jul 2026

Variant characterization in the intrinsically disordered human proteome

Proteome-wide prediction and structural modeling of disordered protein interaction interfaces advance characterization of disease-associated variants in disordered protein regions.

D. Hubrich, Jesús Alvarado Valverde, C. Y. Lee et al. · 0 citations
Open access Aug 2026

Large-scale structure prediction of DUF-containing protein-protein interactions

Whether AlphaFold 3 complex prediction, combined with STRING evidence and domain-level analysis of interfaces and interaction partners, can help identify and characterize DUF-containing proteins and suggest roles for DUF4130 in nucleic-acid-associated radical-SAM biology and DUF5819 in a bacterial system related to vitamin-K-dependent carboxylation are suggested.

Lino Riepenhausen, Francesco Costa, Antonina Andreeva et al. · 0 citations
Open access Jul 2026

Interpretable Prediction of Phase Separation and Disease Variant Effects in Intrinsically Disordered Regions

An interpretable ensemble machine-learning framework that integrates protein language model embeddings of sequence and predicted structure to predict LLPS propensity and classify proteins as self-separating or partner-dependent and identifies critical phase-separating regions and quantifies mutation-induced perturbations in LLPS.

Mingjie Zhao, Sushant Kumar · 0 citations
Open access Jul 2026

Benchmarking AI Protein Structure Predictors Reveals a Persistent Bias in Multi-State Proteins

Protein structure predictors achieve high single-state accuracy, but it remains unclear whether they can recover functionally relevant conformational ensembles or account for the presence of ligands and/or binding partners. Here, we benchmark AlphaFold3, Boltz-2, Chai-1, and BioEmu on four canonical multi-state proteins (Pf-MATE, LAO, SecA, and β2AR), quantifying state bias and sampling breadth against experimental reference structures. Models frequently default to a dominant state represented in the PDB; small-molecule ligands have weak or inconsistent effects, while large protein partners drive clear conformational switching between states. Multiple sequence alignment (MSA)-based approaches (AF-Cluster and random subsampling) recapitulate similar biases, indicating that this behavior is not unique to newer architectures. These results underscore current limitations for multi-state protein structure prediction and structure-guided ligand discovery. TOC Graphic

Muhui Ye, Yu-Hong Wang, M. Brogi et al. · 0 citations
Open access Aug 2026

BENDER: A Cross-taxon IDP Simulation Database Reveals Conserved Sequence-Ensemble Laws Across the Tree of Life

Intrinsically disordered proteins and regions are found across all kingdoms of life, yet the computational characterisation of their conformational ensembles has remained almost entirely confined to the human proteome. Whether the physics-based force fields developed on eukaryotic sequences remain reliable for taxonomically distant organisms, and whether the sequence–ensemble relationships they reveal reflect conserved physical laws or the peculiarities of a single evolutionary window, are questions fundamental to the field. Here we introduce BENDER, a dataset of 11,533 IDP sequences spanning 13 taxonomic groups, each simulated under CALVADOS-2 molecular dynamics and annotated with ensemble-level geometric and novel contact-network properties, together with per-sequence pi–pi and cation–pi contact frequencies linked to phase-separation propensity. We show that CALVADOS-2 ensembles agree strongly with an orthogonal structural reference across the full dataset, with both held-out taxa performing above the dataset median, and that direct comparison against a second independently parameterised force field reveals no systematic scaling-exponent bias. We find that cross-taxon training data improves out-of-distribution ensemble prediction in two independent architectures, and that ensemble contact-network global efficiency is accurately predictable from sequence alone on held-out viral sequences. Positive degree assortativity is conserved across all taxonomic groups, suggesting that hub topology in disordered protein contact networks is a conserved physical feature of sequence-encoded disorder rather than an evolutionary contingency.

J. Velasquez, Taseef Rahman · 0 citations
Open access Aug 2026

A novel benchmark dataset for enzyme function prediction reveals the limitations of state-of-the-art models

It is demonstrated that modern EC predictors largely fail to distinguish catalytically incompetent variants from functional enzymes, and it is proposed that integrating structure-aware negative examples into both training and benchmarking is critical for developing functionally robust models in computational enzymology.

João Sartori, Ana Carolina Ramos Guimarães, Lucas de Almeida Machado · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.