Back to feed
Review

AI-driven discovery of multifunctional peptides: From sequence space to therapeutics.

Jul 2026 · Biotechnology Advances · pp. 108987 · 0 citations · 105 references
Medicine

TL;DR

This review systematically examines the key methodological innovations, including peptide representation learning, multi-modal fusion strategies, multi-label learning paradigms, and emerging predictive frameworks empowered by deep neural architectures and ProtLM-based embeddings, and summarizes the practical applications of these models in peptide database mining, functional mechanism interpretation, and mutation effect prediction.

Abstract

Multifunctional peptides (MFPs), defined as single peptide sequences endowed with two or more biological activities, have emerged as a highly promising class of therapeutic candidates for the intervention of complex and multifactorial diseases. Nevertheless, their discovery has long been constrained by the intrinsic limitations of conventional high-throughput experimental screening and heuristic, rule-based computational strategies, both of which are labor-intensive, costly, low-yield, and poorly suited to the vast combinatorial landscape of peptide sequence space. Recent advances in artificial intelligence (AI), particularly deep learning and protein language models (ProtLMs), have fundamentally transformed this landscape by enabling data-driven, scalable, and representation-rich frameworks for MFP identification. These approaches have substantially improved the ability to capture contextual, structural, physicochemical, and evolutionary determinants of peptide multifunctionality, thereby opening new avenues for systematic peptide discovery. Given the rapid and fragmented progress across this field, a review is imperative to provide an integrated understanding of the methodological landscape, practical applications and persistent challenges in AI-driven MFP discovery. In this review, we first systematically examine the key methodological innovations, including peptide representation learning, multi-modal fusion strategies, multi-label learning paradigms, and emerging predictive frameworks empowered by deep neural architectures and ProtLM-based embeddings. We then summarize the practical applications of these models in peptide database mining, functional mechanism interpretation, and mutation effect prediction. Finally, we outline critical challenges and promising directions that remain in class imbalance, incomplete annotation, limited cross-dataset generalizability, and insufficient integration of drug-likeness and developability criteria.

View source

Similar papers

Review Open access Jun 2026

Targeting the Undruggable: Deep Learning-Driven Design of Peptide Therapeutics in Cancer

How advances in artificial intelligence and computational modeling may reshape the rational design of next-generation peptide therapeutics is explored and an integrated experimental–computational framework is proposed to facilitate the development of clinically actionable candidates is proposed.

Ha Thi Ngoc Nguyen, B. Le, Nhung Thi Hong Van et al. · 0 citations
Review Open access Jul 2026

Artificial intelligence catalyzes antimicrobial peptide design

With broad-spectrum, low resistance, and multifunctional properties, antimicrobial peptides (AMPs) are promising therapeutic agents against drug-resistant pathogens, yet their discovery and optimization still remain challenging due to the complexity of sequence-function associations. Artificial intelligence (AI), through the construction of comprehensive data-driven models that assisted with miscellaneous learning strategies, enables de novo peptide design by learning latent representations inherent in peptide sequences as well as their biological properties to ensure physically plausible and biologically relevant predictions. Consequently, this paradigm enhances the likelihood of designing peptide candidates with significantly improved therapeutic potential, reducing resource-intensive trial-and-error processes and revealing the transformative impact of computational innovation in advancing next-generation therapeutics. Here, we provide a snapshot of this field and survey two modes of AI-driven technologies for AMP design, one concentrated on identifying whether current data possess antimicrobial activity (identification-oriented) and the other on generating AMP candidates with potential therapeutic properties (generation-oriented). We also highlight the challenges and limitations that still hinder AMP development even accelerated by AI, as well as the foreseeable prospects, from finer-grained explorations to model-driven data enrichment and model enhancement.

Yongqiang Liu, Jie Hu, Ning Zhang et al. · 0 citations
Open access Aug 2026

A distribution-aware and functionally relevant novel framework for generation and discovery of bioactive peptides

Recent advances in artificial intelligence have accelerated the discovery of bioactive peptides by enabling computational exploration of the vast peptide sequence space. However, existing peptide generation approaches generally rely on either distribution-learning models, which generate biologically realistic sequences but do not consistently optimize functional activity, or optimization-based methods, which maximize prediction confidence while often deviating from the underlying distribution of experimentally validated peptides. To address this limitation, a two-phase generative–evolutionary framework is proposed that integrates distribution learning with evolutionary optimization. In the first phase, Variational Autoencoders (VAE), Autoregressive Transformers (ART), and Token Diffusion Transformers (TDT) are used to generate biologically plausible seed peptides. In the second phase, these peptides were used as initial seed for Hill Climbing optimization procedure that iteratively improves fitness function score. The proposed two-phase framework was evaluated using a dataset of experimentally validated IL-2-inducing peptides. Evaluation using independent IL-2 prediction models showed that Autoregressive Transformer combined with Hill Climbing achieved the best overall performance, achieving the mean IL-2 induction confidence score of 0.96 while reducing KL divergence from 2.26 for standalone Hill Climbing to 0.75. A case study on an independent IL-13 inducing peptide dataset showed similar trends, with ART initialized Hill Climbing achieving the mean IL-13 induction score of 0.99 while reducing KL divergence from 1.76 to 0.59. Overall, the framework provides a generalizable approach for balancing functional optimization and distributional realism and can be applied to peptide discovery and data augmentation in imbalanced biological datasets thereby generating high confidence peptides for wet lab validation. Highlights Proposed a two-phase framework for bioactive peptide generation with potential to address class imbalance in peptide classification tasks. Performed a systematic comparison of distribution-learning and optimization-based approaches for peptide generation. Combined distribution-learning models for sequence generation with optimization algorithms for improving peptide functional properties. Demonstrated the applicability of the proposed framework across multiple bioactive peptide datasets.

Rachit Abhigyan, Vikas Sood, Pooja Arora et al. · 0 citations
Open access Jul 2026

PeptiVerse: A unified platform for therapeutic peptide property prediction

PeptiVerse is a unified platform that leverages large foundation models to predict diverse peptide developability properties from both amino acid sequences and SMILES representations, enabling accessible, scalable analysis for peptide drug design.

Yinuo Zhang, Sophia Tang, Tong Chen et al. · 5 citations · ⚡1
Open access Jul 2026

A Machine Learning Framework for Short Peptide Sequence Optimization

A data-driven, multi-objective peptide design framework that inte-grates sequence-to-feature transformations using Fast Fourier Transform - based representations, and metric-learning based optimization strategies, to provide an interpretable and computationally efficient alternative for peptide design under limited-data constraints.

A. Trinh · 0 citations
Aug 2026

Machine-Learning-Guided Design of Antifreezing Peptides

An unsupervised machine-learning framework that leverages hybrid high-dimensional peptide representations to discover high-performance AFPT families without requiring 3D structures or large labeled data sets is presented and demonstrates how unsupervised hybrid-feature learning can reveal actionable biophysical design rules from sequence data alone.

Nazmul Shuzan, Jialun Wei, Jie Zheng · 0 citations