Back to feed

Machine-Learning-Guided Design of Antifreezing Peptides

Aug 2026 · Journal of Chemical Information and Modeling · 0 citations · 69 references

TL;DR

An unsupervised machine-learning framework that leverages hybrid high-dimensional peptide representations to discover high-performance AFPT families without requiring 3D structures or large labeled data sets is presented and demonstrates how unsupervised hybrid-feature learning can reveal actionable biophysical design rules from sequence data alone.

Abstract

Antifreeze peptides (AFPTs) offer a potentially nontoxic, sequence-programmable alternative to conventional cryoprotectants for preserving biological materials, yet poorly defined sequence–activity relationships continue to limit rational design. Natural AFPTs are often weak, scarce, or context-dependent, and existing design strategies rely on incremental motif tuning with low hit rates and limited interpretability. Here, we present an unsupervised machine-learning framework that leverages hybrid high-dimensional peptide representations to discover high-performance AFPT families without requiring 3D structures or large labeled data sets. We curated the largest annotated AFPT benchmark to date (n = 719) and embedded sequences in a feature space combining physicochemical descriptors with protein language model (PLM) embeddings. Unsupervised clustering resolved distinct active families within the sequence landscape, quantitatively validated by a subset of 107 peptides with measured single-crystal ice-growth rates. Mechanistic interrogation uncovered a dual-signal architecture not previously codified at the peptide level: (i) a regularly spaced polar ice-binding face encoded by primary-sequence motifs, coupled with (ii) a rigid, glycine-depleted scaffold captured only by latent PLM features. A logistic regression classifier trained on the minimal 10-feature set achieved near-perfect separability of active versus inactive families (AUC = 0.98), confirming the generality of the dual-signal rule. Guided by this interpretable blueprint, we designed 14 de novo peptides─10 dual-signal positive designs and 4 negative controls─and validated them alongside 3 literature-reported benchmarks through multiple orthogonal assays. As predicted, negative controls showed minimal activity across all metrics, whereas dual-signal designs exhibited strong ice recrystallization inhibition (IRI activity up to ∼55%), substantial temperature (down to −3.45 °C), and high red-blood-cell post-thaw recovery (85–92%). This work establishes a generalizable, presynthesis prioritization framework for antifreeze peptide engineering and demonstrates how unsupervised hybrid-feature learning can reveal actionable biophysical design rules from sequence data alone.

View source

Similar papers

Open access Jun 2026

Computational Redesign of an Antifreeze Protein Using Deep Learning

Antifreeze proteins (AFPs) found in various cold-adapted organisms inhibit ice growth and are of interest for applications in food products, cryopreservation, agriculture, and materials science. Although high-resolution structures are available for several AFPs, the amino acids required for full antifreeze activity remain incompletely defined, and the development of AFP variants with properties such as enhanced solubility, high expression yield, and improved thermostability may further facilitate applications. Here, we used the deep learning model ProteinMPNN to redesign the globular fish antifreeze protein AFPIII, keeping the previously reported ice-binding residues fixed. We readily obtained sequences confidently predicted to adopt AFPIII’s structure and we selected five designed variants for expression, all of which expressed efficiently in E. coli. Circular dichroism spectroscopy showed that two of these variants retained secondary structure elements consistent with AFPIII, whereas the other three exhibited structural differences. One design was predicted and experimentally confirmed to have increased thermostability. All five variants displayed measurable thermal hysteresis activity. However, none reached the activity of wild-type AFPIII, suggesting that maintaining the currently established set of ice-binding residues is not sufficient to fully preserve this AFP’s function; other, unidentified residues can also impact its activity. Our findings highlight the value of deep learning-based protein design methods both for generating AFP variants with desirable properties and for uncovering gaps in existing knowledge of well-characterized AFPs.

Cianna N. Calia, Arthur J. Altunc, Rosemary J. Eufemio et al. · 0 citations
Review Jul 2026

AI-driven discovery of multifunctional peptides: From sequence space to therapeutics.

This review systematically examines the key methodological innovations, including peptide representation learning, multi-modal fusion strategies, multi-label learning paradigms, and emerging predictive frameworks empowered by deep neural architectures and ProtLM-based embeddings, and summarizes the practical applications of these models in peptide database mining, functional mechanism interpretation, and mutation effect prediction.

Zhiqiang Liang, Yupeng Hao, Junjie Chen · 0 citations
Review Open access Jun 2026

Targeting the Undruggable: Deep Learning-Driven Design of Peptide Therapeutics in Cancer

How advances in artificial intelligence and computational modeling may reshape the rational design of next-generation peptide therapeutics is explored and an integrated experimental–computational framework is proposed to facilitate the development of clinically actionable candidates is proposed.

Ha Thi Ngoc Nguyen, B. Le, Nhung Thi Hong Van et al. · 0 citations
Open access Jul 2026

AI-guided discovery for low-resource peptide engineering using evolutionary scale modeling

Reliable estimation of downstream performance in low-data peptide machine learning is critical for guiding early-stage AI-driven peptide engineering. Yet, it is often unclear how to assess whether a model will be effective in iterative discovery settings. Here, we show that the cross validation R² score can serve as a simple and robust proxy for predicting active learning workflow performance, enabling early-stage evaluation of model suitability for sequential peptide optimization. To support this, we introduce SCARSE, a machine learning framework combining ESM-2 protein language model embeddings with Gaussian process regression and extremely randomized trees classification, designed for low-resource peptide property prediction (20–500 training samples). We benchmark SCARSE across 23 peptide and small-protein datasets covering substitution and indel variants, antimicrobial peptides, cell-penetrating peptides, and toxic/non-toxic peptides. SCARSE significantly outperforms a hand-engineered descriptor baseline on substitution and indel tasks, while comparable performance was achieved on shorter peptide non-mutant datasets where simpler descriptors capture enough of the signal. In simulated active learning workflows, SCARSE consistently outperforms baseline and random sampling strategies. Notably, we demonstrate that CV R² computed from as few as 50 labeled peptides can be sufficient to estimate final active learning end-point performance, providing a practical, data-efficient criterion for deciding whether a given dataset combined with SCARSE is suitable for iterative peptide discovery. SCARSE is released as a pip package and is available via HuggingFace Spaces to facilitate integration into peptide engineering workflows.

Leo Andrekson, Robin Rydbergh, Rocío Mercado et al. · 0 citations
Open access Dec 2025

DeepAden: an explainable machine learning method for predicting the substrate specificity of nonribosomal peptide synthetases

Microbial non-ribosomal peptides (NRPs) exhibit remarkable structural diversity and serve as valuable sources of lead compounds for clinical drug development. The biosynthesis of NRPs relies on non-ribosomal peptide synthetases (NRPSs), in which adenylation (A) domains play a pivotal role in defining the core structure by selectively recognizing and activating amino acid substrates. Accurately predicting the substrate specificities of A-domains is thus essential for understanding the core structural and biosynthetic logic of NRPs. Here, we present DeepAden, a two-stage deep learning framework. In the first stage, a graph attention network (GAT)-based model localizes 27-residue binding pockets within 6 Å of bound substrates and convert these into pocket representations. In the second stage, pocket representations are then encoded alongside substrate information using pretrained language models, and aligned using contrastive learning. In addition, we introduce a SHapley Additive exPlanations (SHAP)-guided data augmentation strategy to mitigate class imbalance and improve robustness, particularly for nonproteinogenic substrates. DeepAden achieves competitive performance compared with state-of-the-art tools on a benchmark dataset, and enabled the identification of two Streptomyces NRPS gene clusters through accurate A-domain substrates specificity predictions. DeepAden offers a powerful tool for precise pocket localization and robust substrate prediction, accelerating the discovery and characterization of novel NRP natural products for future work. The DeepAden web server is available at https://deepnp.site/.

Jiaquan Huang, Liangjun Ge, Yaxin Wu et al. · 0 citations
Open access Jul 2026

A Machine Learning Framework for Short Peptide Sequence Optimization

A data-driven, multi-objective peptide design framework that inte-grates sequence-to-feature transformations using Fast Fourier Transform - based representations, and metric-learning based optimization strategies, to provide an interpretable and computationally efficient alternative for peptide design under limited-data constraints.

A. Trinh · 0 citations