Skip to content
Open access

MAXWELL: Calibrating the probabilistic outputs of protein language models to the mutation-induced stability change landscape

Aug 2026 · bioRxiv · 0 citations · 61 references
Biology

TL;DR

MAXWELL (Matrix-wise Landscape Learning), a novel post-training method that calibrates the probabilistic outputs learned by protein language models during pretraining to generate mutational landscapes that quantify the effects of individual amino acid substitutions on protein stability, is introduced.

Abstract

Designing mutations that enhance protein stability is a central goal in protein engineering. However, experimentally screening large numbers of candidate mutations is costly and time-consuming, creating a strong need for computational methods that can identify potentially stabilizing mutations. Among these approaches, protein language models are particularly promising because they learn context-dependent amino acid preferences from large-scale sequence and structure datasets. Nevertheless, most existing stability prediction methods use these models primarily as feature extractors and do not fully exploit the amino acid probability distributions they encode. Here, we introduce MAXWELL (Matrix-wise Landscape Learning), a novel post-training method that calibrates the probabilistic outputs learned by protein language models during pretraining to generate mutational landscapes that quantify the effects of individual amino acid substitutions on protein stability. When applied to ProteinMPNN, MAXWELL yields a state-of-the-art predictor of the effects of protein mutations on stability, outperforming ThermoMPNN and other representative methods on a curated benchmark of experimentally measured stability changes. We next applied MAXWELL to the design of ten single-point mutations in the DhaA dehalogenase, seven of which (70%) increased thermal stability. Among them, G171W showed the largest improvement, with a measured ΔTm of 4.91 °C. These experimental results establish MAXWELL as a novel post-training strategy for protein language models and a practical framework for designing stabilizing mutations. Repository https://github.com/ai4protein/Venus-MAXWELL

Read PDF

Similar papers

Open access Jul 2026

Protein language models learn underlying mutation biases alongside fitness landscapes

A parameter sweep is implemented to explicitly couple empirical nucleotide mutational supply from PLM-assessed amino acid substitution pseudo-probabilities across evolutionary forecasting tasks and finds that base PLMs implicitly learn generic nucleotide-level mutational constraints, an effect strongly amplified by virus-specific fine-tuning.

O. MacLean, Kieran D. Lamb, Spyros Lytras et al. · 0 citations
Open access Aug 2026

A unified predictor of protein stability changes across all mutation types via implicit structure learning

UniStab is introduced, an end-to-end framework for predicting stability changes across all mutation types by leveraging the implicit geometric reasoning of a pre-trained folding model and demonstrates state-of-the-art performance, particularly in the challenging scenarios of multi-point mutations and indels.

Hong Tan, Sheng-Geng Lin, Yi Xiong · 0 citations
Open access Aug 2026

An ensemble learning framework for protein stability prediction with enhanced recognition of stabilizing mutations

Three modeling frameworks are developed, including models based on handcrafted features, models using embedding representations extracted from ProteinMPNN, and ensemble models integrating a diverse set of state‐of‐the‐art predictors integrating a diverse set of state‐of‐the‐art predictors.

Yang Liu, Jian Zhang, Minghui Li · 0 citations
Open access Aug 2026

Benchmarking Deep Learning Predictions of Mutation-Induced Fold Switching

A systematic NMR-characterized dataset of mutants of the GA/GB model fold-switching system is presented and it is found that this benchmark revealed variable and position-dependent performance across methods, with certain AlphaFold2-based algorithms able to predict mutant effects at individual sites, indicating some understanding of physical effects of residue substitutions.

Nathaniel R. Felbinger, K. Carillo, Yihong Chen et al. · 0 citations
Open access Aug 2026

Aligning protein-generative models to experimental fitness with ProteinDPO

This work demonstrates how to provide task-specific information without losing the general knowledge learned during pretraining by using direct preference optimization to align a structure-conditioned protein language model to preferentially generate stable protein sequences.

Talal Widatalla, Ashir Borah, Samuel H. King et al. · 1 citation
Open access Aug 2026

Accelerating protein engineering: an integrated framework combines protein language models and epistatic landscape modeling

MULTI-evolve is a model guided, universal, targeted installation of multimutants framework that rapidly designs hyperactive multimutant proteins and improves the identi fi cation of productive mutations compared with individual PLMs alone.

J. Koo, Young-Ho Park, Sun-Uk Kim · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.