Skip to content
Open access

Artificial intelligence-driven prediction and design of cell-penetrating peptides for advanced drug delivery system

Aug 2026 · Frontiers in Pharmacology · Vol 17 · 0 citations · 59 references
Medicine

TL;DR

An artificial intelligence-driven framework that integrates biochemical rules derived from large language models (LLMs) with conventional peptide descriptors for CPP prediction and design is developed and demonstrates that combining LLM-derived biochemical knowledge with machine learning improves interpretable CPP prediction and candidate prioritisation.

Abstract

Background Cell-penetrating peptides (CPPs) are promising delivery vectors for transporting therapeutic agents across cellular membranes. However, their rational design remains challenging because the relationship between peptide sequence and translocation efficiency is highly complex and nonlinear. Methods In this study, we developed an artificial intelligence-driven framework that integrates biochemical rules derived from large language models (LLMs) with conventional peptide descriptors for CPP prediction and design. Interpretable rules extracted from GPT-4o and DeepSeek were encoded as binary feature vectors and combined with sequence-based descriptors to construct hybrid machine learning models. Model performance was evaluated on the benchmark CPP924 dataset using repeated stratified cross-validation, and the optimized models were further used for de novo CPP generation. The resulting candidates were subsequently assessed using multiple established computational benchmarks. Results The top-performing hybrid classifier achieved a cross-validated accuracy of 0.91 ± 0.03 (best single held-out split, 0.94) on the CPP924 dataset. The LLM-derived rules outperformed conventional physicochemical and fingerprint descriptors and matched amino-acid composition; integrating the rule and composition features yielded the best overall classifier. Using the optimized RF-GPT-Fre and RF-DS-Fre models, we generated six de novo CPP candidates that are sequence-novel (≤53% identity to any training peptide) and retain CPP-like composition and structural features. In-silico evaluation across established tools supports computational prioritisation of these candidates for experimental testing. Conclusion These findings demonstrate that combining LLM-derived biochemical knowledge with machine learning improves interpretable CPP prediction and candidate prioritisation. This study provides a reproducible computational strategy for peptide engineering and establishes a basis for the experimental evaluation of next-generation drug-delivery vehicles.

Read PDF

Similar papers

Review Jul 2026

AI-driven discovery of multifunctional peptides: From sequence space to therapeutics.

This review systematically examines the key methodological innovations, including peptide representation learning, multi-modal fusion strategies, multi-label learning paradigms, and emerging predictive frameworks empowered by deep neural architectures and ProtLM-based embeddings, and summarizes the practical applications of these models in peptide database mining, functional mechanism interpretation, and mutation effect prediction.

Zhiqiang Liang, Yupeng Hao, Junjie Chen · 0 citations
Review Open access Jul 2026

Artificial intelligence catalyzes antimicrobial peptide design

A snapshot of AI-driven technologies for AMP design is provided and two modes of AI-driven technologies for AMP design are surveyed, one concentrated on identifying whether current data possess antimicrobial activity and the other on generating AMP candidates with potential therapeutic properties (generation-oriented).

Yong-Qiang Liu, Jie Hu, Ning Zhang et al. · 0 citations
Review Sep 2026

Computational Platforms for Membrane-Active Peptide Therapeutic Design

Peptide-based therapeutics are gaining popularity as next-generation drugs because of their potent bioactivity, high specificity, and broad applications in immunomodulatory, antiviral, anticancer, and antimicrobial therapy. Membrane-active peptides (MAPs) are particularly promising, yet poor solubility, stability, pharmacokinetics, manufacturing complexity, and limited delivery options have restricted their clinical translation. This review focuses on computational approaches developed to address these challenges at the peptide-design stage and bridge the gap between preclinical promise and clinical utility. The goal here is to provide a broad overview of tools for in silico MAP design, prediction of functional structural ensembles and physicochemical properties, structure–function relationships, drug-target interactions, and formulation strategies for in vivo delivery. These approaches span bioinformatics, molecular dynamics and molecular modeling, and emerging artificial intelligence and machine learning platforms, including hybrid computational-experimental strategies, chemical modification and conjugation, and nanotechnology-based delivery systems. These tools are improving key peptide–drug properties, including plasma stability, target specificity, efficacy, and delivery. Continued advances in computational design may enable MAPs to address challenging diseases, engage traditionally undruggable targets, and translate their substantial therapeutic potential into clinically effective treatments.

Unknown authors · 0 citations
Aug 2026

Systematic Benchmarking of AI-Based Molecular Generation Models for Structure-Based Drug Design

A state-aware functional classifier (SAFC) is developed that integrates molecular dynamics derived receptor ensembles, ensemble docking and protein ligand interaction graphs that provides dynamics-aware functional activity rankings for generated molecules that were partly complementary to docking, drug-likeness and synthetic accessibility scores.

H. Kumar, Zheng-Xiao Yang, Yankai Yu et al. · 0 citations
Open access Aug 2026

Knowledge-driven multimodal mutual learning for cell line-targeted anticancer peptide prediction.

MOTIVATION As unique drugs positioned between small and macro molecules, anticancer peptides (ACPs) hold great potential in oncotherapy owing to their high selectivity and low toxicity. Nowadays, computational ACP prediction has emerged as a cost-effective alternative to bioassay screening, but most methods are limited to identifying bioactivity and fail to resolve tumor cell-specific targeting, primarily because of the sparse annotated data. RESULTS To fill this gap, we integrate a hybrid dataset compiled from five well-established peptide databases and propose TargetPC, a deep learning method tailored for cell line-targeted ACP prediction. TargetPC encodes multimodal representations of ACPs and cell lines via pretrained protein and omics models, and combines them via hierarchical intra- and inter-modal fusion for targeting prediction. This combination is further augmented by a mutual learning paradigm that distills domain knowledge from both ACP and cell line, enabling improved generalization under sparse supervision. Experimental results on the hybrid dataset demonstrate the effectiveness of TargetPC, which outperforms the state-of-the-art baselines in terms of prediction accuracy, and maintains strong generalization to unseen ACPs and cell lines. When extended to out-of-distribution samples, TargetPC has successfully screened dozens of novel ACPs targeted to breast cancer cells and uncovered biological motifs underlying its predictions. As a result, our TargetPC is expected to serve as a versatile tool for lead ACP discovery at a lower burden. AVAILABILITY The source code and data are available at GitHub (https://github.com/liuxuan666/TargetPC).

Xuan Liu, Jian Zhang, Chong-Yang Chen et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.