Back to feed
Open access

A Machine Learning Framework for Short Peptide Sequence Optimization

Jul 2026 · Proceedings of the 3rd Foundations of Process/Product Analytics and Machine Learning (FOPAM 2026) · 0 citations · 1 references

TL;DR

A data-driven, multi-objective peptide design framework that inte-grates sequence-to-feature transformations using Fast Fourier Transform - based representations, and metric-learning based optimization strategies, to provide an interpretable and computationally efficient alternative for peptide design under limited-data constraints.

Abstract

Designing peptides plays an important role in applications ranging from therapeuticsand biomaterials to diagnostics. However, due to the large combinatorial sequence spaceand the high cost and time required for experimental screening, experimental trial anderror approaches are prohibitively expensive. Furthermore, peptide design is inherentlya multi-objective problem that requires simultaneous optimization of different propertiessuch as biological activity, stability, solubility, and safety. These challenges motivate theuse of computational design strategies. Traditional physics-based and sequence-alignmentmethods often struggle to handle variable length sequences and often rely on structuralinformation that is unavailable for many peptides.1, 2 More recently, deep learning modelssuch as AlphaFold, ESM, and diffusion-based approaches have transformed protein mod-eling3, 4, 5 . However, their large data requirements, high computational cost, and black-boxnature reduce their practicality for deterministic multi-objective optimization in limited-data settings.6This work proposes a data-driven, multi-objective peptide design framework that inte-grates sequence-to-feature transformations using Fast Fourier Transform (FFT) - basedrepresentations,7, 8 interpretable feature attribution through GroupSHAPLEY, and metric-learning based optimization strategies. A bidirectional mapping between sequence spaceand feature space is introduced to identify critical feature contributions and improve inter-pretability during peptide optimization.The primary focus of this poster is the optimization component of the framework. Specif-ically, distance metric learning methods, including Neighborhood Component Analysis(NCA)9 and Large Margin Nearest Neighbor (LMNN),10 are investigated to maximizeclass separation between peptide groups and identify discriminative feature representations.Comparative analyses were performed to evaluate the robustness, tunability, and optimiza-tion behavior of these methods on both synthetic and peptide-representative datasets. UsingSupport Vector Machine (SVM) as the predictive model, the proposed methods demonstrateimproved optimization efficiency through dimensionality reduction in synthetic data exper-iments. Multiple optimization constraints can be incorporated to assess the robustness,scalability, and adaptability of the framework across different design scenarios. The pro-posed framework aims to provide an interpretable and computationally efficient alternativefor peptide design under limited-data constraints.

Read PDF

Similar papers

Review Open access Jun 2026

Targeting the Undruggable: Deep Learning-Driven Design of Peptide Therapeutics in Cancer

How advances in artificial intelligence and computational modeling may reshape the rational design of next-generation peptide therapeutics is explored and an integrated experimental–computational framework is proposed to facilitate the development of clinically actionable candidates is proposed.

Ha Thi Ngoc Nguyen, B. Le, Nhung Thi Hong Van et al. · 0 citations
Review Jul 2026

AI-driven discovery of multifunctional peptides: From sequence space to therapeutics.

This review systematically examines the key methodological innovations, including peptide representation learning, multi-modal fusion strategies, multi-label learning paradigms, and emerging predictive frameworks empowered by deep neural architectures and ProtLM-based embeddings, and summarizes the practical applications of these models in peptide database mining, functional mechanism interpretation, and mutation effect prediction.

Zhiqiang Liang, Yupeng Hao, Junjie Chen · 0 citations
Open access Jul 2026

AI-guided discovery for low-resource peptide engineering using evolutionary scale modeling

Reliable estimation of downstream performance in low-data peptide machine learning is critical for guiding early-stage AI-driven peptide engineering. Yet, it is often unclear how to assess whether a model will be effective in iterative discovery settings. Here, we show that the cross validation R² score can serve as a simple and robust proxy for predicting active learning workflow performance, enabling early-stage evaluation of model suitability for sequential peptide optimization. To support this, we introduce SCARSE, a machine learning framework combining ESM-2 protein language model embeddings with Gaussian process regression and extremely randomized trees classification, designed for low-resource peptide property prediction (20–500 training samples). We benchmark SCARSE across 23 peptide and small-protein datasets covering substitution and indel variants, antimicrobial peptides, cell-penetrating peptides, and toxic/non-toxic peptides. SCARSE significantly outperforms a hand-engineered descriptor baseline on substitution and indel tasks, while comparable performance was achieved on shorter peptide non-mutant datasets where simpler descriptors capture enough of the signal. In simulated active learning workflows, SCARSE consistently outperforms baseline and random sampling strategies. Notably, we demonstrate that CV R² computed from as few as 50 labeled peptides can be sufficient to estimate final active learning end-point performance, providing a practical, data-efficient criterion for deciding whether a given dataset combined with SCARSE is suitable for iterative peptide discovery. SCARSE is released as a pip package and is available via HuggingFace Spaces to facilitate integration into peptide engineering workflows.

Leo Andrekson, Robin Rydbergh, Rocío Mercado et al. · 0 citations
Aug 2026

Machine-Learning-Guided Design of Antifreezing Peptides

An unsupervised machine-learning framework that leverages hybrid high-dimensional peptide representations to discover high-performance AFPT families without requiring 3D structures or large labeled data sets is presented and demonstrates how unsupervised hybrid-feature learning can reveal actionable biophysical design rules from sequence data alone.

Nazmul Shuzan, Jialun Wei, Jie Zheng · 0 citations
Jul 2026

PepOSX-AI: CPP - an interpretable transformer-based deep learning model for prediction of cell-penetrating peptides.

BACKGROUND Cell-penetrating peptides (CPPs) are short-chain molecules capable of enhancing the transmembrane delivery of bioactive substances, displaying extensive application potential in the delivery of functional components and the improvement of their bioavailability. Traditional CPP discovery methods, however, rely on a tedious, step-by-step screening process involving cell and animal experiments, which is highly inefficient. METHODS The deep-learning model, PepOSX-AI: CPP, was developed based on the Transformer architecture and integrative features of six physicochemical descriptors. This involved a series of explorations of model parameters and automated hyperparameter optimization. The model achieved a high area under the curve (AUC) of 0.914 on the dataset. Furthermore, the model effectively captured long-range dependencies in amino acid sequences through a self-attention mechanism, enabling the interpretability of attention patterns. This capability facilitates the identification of key sequence features influencing peptide penetration ability. SIGNIFICANCE AND NOVELTY In comparison to some existing CPP prediction models, PepOSX-AI: CPP demonstrated an accuracy of 91.00% on the application test set, highlighting its acceptable predictive performance. It provides a novel computational tool and theoretical basis for the screening of CPPs. Based on these findings, an online platform was also developed to facilitate user application.

Haowen Chen, Weiwei He, Kaiyan Feng et al. · 0 citations
Open access Aug 2026

A distribution-aware and functionally relevant novel framework for generation and discovery of bioactive peptides

Recent advances in artificial intelligence have accelerated the discovery of bioactive peptides by enabling computational exploration of the vast peptide sequence space. However, existing peptide generation approaches generally rely on either distribution-learning models, which generate biologically realistic sequences but do not consistently optimize functional activity, or optimization-based methods, which maximize prediction confidence while often deviating from the underlying distribution of experimentally validated peptides. To address this limitation, a two-phase generative–evolutionary framework is proposed that integrates distribution learning with evolutionary optimization. In the first phase, Variational Autoencoders (VAE), Autoregressive Transformers (ART), and Token Diffusion Transformers (TDT) are used to generate biologically plausible seed peptides. In the second phase, these peptides were used as initial seed for Hill Climbing optimization procedure that iteratively improves fitness function score. The proposed two-phase framework was evaluated using a dataset of experimentally validated IL-2-inducing peptides. Evaluation using independent IL-2 prediction models showed that Autoregressive Transformer combined with Hill Climbing achieved the best overall performance, achieving the mean IL-2 induction confidence score of 0.96 while reducing KL divergence from 2.26 for standalone Hill Climbing to 0.75. A case study on an independent IL-13 inducing peptide dataset showed similar trends, with ART initialized Hill Climbing achieving the mean IL-13 induction score of 0.99 while reducing KL divergence from 1.76 to 0.59. Overall, the framework provides a generalizable approach for balancing functional optimization and distributional realism and can be applied to peptide discovery and data augmentation in imbalanced biological datasets thereby generating high confidence peptides for wet lab validation. Highlights Proposed a two-phase framework for bioactive peptide generation with potential to address class imbalance in peptide classification tasks. Performed a systematic comparison of distribution-learning and optimization-based approaches for peptide generation. Combined distribution-learning models for sequence generation with optimization algorithms for improving peptide functional properties. Demonstrated the applicability of the proposed framework across multiple bioactive peptide datasets.

Rachit Abhigyan, Vikas Sood, Pooja Arora et al. · 0 citations