Aug 2026· BMC Biology· Vol 24· 0 citations· 64 references
Medicine
TL;DR
TPpred-PepPA is developed, a two-stage hierarchical deep learning framework based on the ProtT5 pre-trained large language model that achieves state-of-the-art predictive performance and provides valuable interpretability for the discovery of multi-functional therapeutic peptides.
Abstract
Therapeutic peptides exert pivotal effects in diverse biological processes, and have attracted significant interest in the field of biomedicine in recent years. However, most existing methods often fail to adequately capture the intricate interactions among amino acid residues and the contextual dependencies within peptide sequences, which hampers the extraction of deep semantic representations and ultimately restricts predictive performance. Moreover, the task of multi-functional therapeutic peptide prediction is inherently constrained by the challenge of imbalanced multi-label classification resulting from long-tailed distribution patterns. In this study, we propose a two-stage hierarchical deep learning framework, named TPpred-PepPA, for the prediction of multi-functional therapeutic peptides based on pragmatic analysis. Specifically, ProtT5 is employed to extract deep semantic representations that capture residue-level contextual information. In the first stage, a transformer-based network is utilized to perform shared representation learning, wherein the encoder model captures the intricate inter-residue interaction to characterize the contextual semantics of peptide sequences. In the second stage, the framework is fine-tuned by incorporating task-specific classifiers and optimizing the classification decision with Asymmetric Loss. Then the dynamic thresholding strategy is utilized to address the long-tail distribution problem, enabling more accurate prediction performance of multi-functional therapeutic peptide. Moreover, we adopted the SHAP analysis and motif identification to interpret feature contributions and identify key functional peptide fragments, respectively. Our experimental results indicate that TPpred-PepPA significantly outperforms all current baseline methods in identifying multi-functional therapeutic peptides and exhibits robust performance in recognizing rare functional categories. We developed TPpred-PepPA, a two-stage hierarchical deep learning framework based on the ProtT5 pre-trained large language model. Compared with existing methods, TPpred-PepPA achieves state-of-the-art predictive performance and provides valuable interpretability for the discovery of multi-functional therapeutic peptides. Finally, a web server has been established and is accessible at http://bliulab.net/TPpred-PepPA.
It is shown that single-sequence PLMs can perform in-context peptide learning without gradient updates, task-specific retraining, or architectural modification, and MPEP conditioning is established as a lightweight strategy for low-data peptide classification.
Joshua Almonte, Minh N. Vu, Andrew Ahn et al.· bioRxiv· 0 citations
This paper proposes TextDTI, a multimodal framework that simultaneously exploits sequential and structural representations and enhances feature alignment through adversarial learning and contrastive loss, resulting in robust and high-performance DTI prediction.
Jiaqi Deng, Senyu Tang, Ji-Jun Tang et al.· Journal of Chemical Informat...· 0 citations
Accurate computational prediction of antiviral peptides (AVPs) can accelerate peptide screening and reduce experimental costs. However, existing deep learning-based methods still suffer from severe class imbalance, over-reliance on handcrafted features and limited interpretability. Here, we propose PMAVP, a multi-task learning framework that integrates the ProtT5 pre-trained protein language model with a Mamba-inspired module for AVP identification and functional activity prediction. We use ProtT5 to extract deep semantic representations from peptide sequences and a Mamba module to capture long-range dependencies at a lower computational complexity. We introduce Focal Loss to mitigate class imbalance and leverage transfer learning to enhance performance on functional activity prediction. Experimental results demonstrate that our model achieves superior performance in terms of prediction accuracy, stability, and computational efficiency. Furthermore, DeepSHAP-based interpretability analysis reveals that the first 40 amino acid residues contribute substantially to AVP prediction.
Pei-Wei Wei, Wei-Hao Su, Qing-Song Qin et al.· International Journal of Mol...· 0 citations
Abstract Motivation Peptides serve as critical mediators in biological systems, regulating essential processes ranging from neurotransmission to immune response. However, their nonlinear sequence–function relationships and immense chemical diversity pose significant challenges for efficient experimental characterization and therapeutic development. While Protein Language Models have advanced biological sequence understanding, they predominantly capture global evolutionary features of full-length proteins, often overlooking the local physicochemical dependencies and short-range residue interactions that are essential for defining peptide bioactivity. Results To address this gap, we leverage the linguistic competence of general-purpose Large Language Models (LLMs) to treat amino acid sequences as “biological text,” bridging natural language supervision with biochemical sequence modeling without relying on explicit structural or evolutionary priors. We curate a peptide-specific instruction dataset, Pep-Instructions, spanning function description, sequence design, property prediction, and physicochemical optimization, and adapt a general-purpose LLM through parameter-efficient instruction tuning. Extensive benchmarking against general-purpose language models shows consistent improvements across the evaluated peptide-centric tasks. In particular, the instruction-tuned model produces more semantically faithful functional descriptions, generates peptide sequences with stronger sequence-level similarity, with representative ESMFold case studies suggesting backbone-level consistency, improves prediction of diverse peptide properties, and enables more reliable directional optimization of physicochemical properties under the adopted in silico evaluation protocols. Overall, these results establish Pep-Instructions as a unified benchmark and demonstrate the value of peptide-specific instruction tuning for peptide understanding, prediction, and design. Availability and implementation Source code and Pep-Instructions are available at https://github.com/kjY7836/pepinstruction. Fine-tuned model weights are available at https://huggingface.co/Codelife176/Pep-instruction.
Kai Yang, Tian-Xiang Wu, Wenbo Zhang et al.· Bioinformatics· 0 citations
Predicting peptide retention time (RT) remains a significant challenge, particularly when training data is limited. In this study, we present MetaRT, a stacked-ensemble machine learning framework designed to predict the RTs from small dataset of hydrophobic peptides. Peptides composed of hydrophobic amino acids—phenylalanine (F), isoleucine (I), methionine (M), and tryptophan (W) were synthesized, and their experimental RTs were measured from the mixture entities. MetaRT utilizes a graph convolutional network (GCN) to extract structural features from the peptide sequences. The MetaRT model architecture employed multiple base learners, integrating the outputs through a meta-learner optimized via hyperparameter tuning and 3-fold cross-validation. Besides, the performance of MetaRT was compared to three ensemble methods - weight averaging, bagging, and boosting. The results demonstrated that structure-based MetaRT outperformed both base learners and the ensemble models, achieving a lower root mean square error (RMSE) of 0.08 and a maximum RT deviation of approximately 1.4 min. Compared to the prediction performance on molecular descriptors inclusion, the structure-guided model consistently performed well in terms of RMSE. Notably, MetaRT accurately predicted the RTs of sequence isomers by leveraging the structural features, with deviations ranging from 0.2 to 1.3 min. In contrast, descriptor-based model showed increased prediction error for the isomeric sequences. For peptides with lower hydrophobicity that were not included in the training data, the structure-based predictions led to the maximum deviation of 4.9 min from the experimental RTs. The entire predicted RTs were subsequently validated by linear regression analyses with the corresponding experimental values. These findings highlight the potential of MetaRT as a structure-based predictive tool for improving RT prediction accuracy, especially in data-limited scenarios. Future work will focus on enhancing the robustness of MetaRT by incorporating a wider variety of peptide classes to further refine its predictive capabilities.
R. A. Mahmood, M. Mahin, R. Alam et al.· Journal of Analytical Scienc...· 0 citations
Overall, DrugPLMFormer provides a reproducible, leakage-aware framework for retrospective sequence-based druggability screening and target prioritization, while prospective validation and experimental confirmation remain necessary before operational deployment.
Z. Kafi, Khosro Rezaee, Hossein Eslami· Journal of King Saud Univers...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.