A new concept termed DeepPepQSAR that integrates deep learning into traditional PepQSAR is proposed to quantitatively model, predict, and interpret the BAP universe in an all-in-one manner for artificial intelligence (AI)-driven big-data BAP discovery.
This review systematically examines the key methodological innovations, including peptide representation learning, multi-modal fusion strategies, multi-label learning paradigms, and emerging predictive frameworks empowered by deep neural architectures and ProtLM-based embeddings, and summarizes the practical applications of these models in peptide database mining, functional mechanism interpretation, and mutation effect prediction.
The widespread misuse of antibiotics has led to a global antimicrobial resistance crisis, highlighting the urgent need for novel antibacterial strategies. Short antimicrobial peptides (sAMPs), while maintaining strong antimicrobial activity, offer superior bioavailability and synthetic feasibility, thus holding great promise in the development of next-generation antibiotics. In recent years, AI-based approaches have achieved notable progress in AMP prediction; however, most existing models are trained primarily on medium and long peptides, resulting in limited accuracy and representation capability when applied to identify sAMPs. To address this problem, a novel prediction model, PPsAMP, is proposed in this paper. First, the protein language model ProtBert-BFD is fine-tuned by sAMPs and non-sAMPs to extract more discriminative representations, which are then integrated with physicochemical features through a cross-attention mechanism. The fused representation is further processed by a feature learning module to achieve the identification of sAMP. The feature learning module consists of a multihead self-attention mechanism and feedforward layers, with residual connections added to enhance generalization ability. Experimental results demonstrate that PPsAMP significantly outperforms state-of-the-art models for identifying sAMPs. Moreover, PPsAMP has identified 14,839 candidate sAMPs from environmental metagenomes, most of which have not been previously reported. The predicted MIC values indicate that they possess potential antibacterial activity. PPsAMP is freely available at https://github.com/shengxiliu/PPsAMP.
Shengxi Liu, Xizhe Gao, Jingyu Wang et al.· Journal of Chemical Informat...· 0 citations
The proposed two-phase generative–evolutionary framework provides a generalizable approach for balancing functional optimization and distributional realism and can be applied to peptide discovery and data augmentation in imbalanced biological datasets thereby generating high confidence peptides for wet lab validation.
Background: Orientothele washanensis is a venomous spider with considerable ecological and scientific importance. Its venom, characterized by complex composition and ease of collection, serves as a valuable resource for the discovery of natural peptide drugs. Conventional wet-lab screening methods are limited by rigorous experimental conditions, high resource consumption, and long research cycles, which hinder the efficient identification of functional peptides from the venom gland transcriptome of this spider. Methods: To address these technical bottlenecks, this study developed a novel deep learning model named PepPI-DRN for peptide-protein interaction prediction. The model integrated a residual equivariant graph neural network, a residual 1-dimensional convolutional neural network, and a dual-modal attention mechanism by leveraging both sequence and structural features of peptides and proteins. Results: Results on an independent test set indicated that PepPI-DRN achieved the competitive or superior performance compared with state-of-the-art methods on multiple key evaluation metrics. Candidate peptides with high interaction probabilities against targets were obtained from the venom gland transcriptome of Orientothele washanensis. Furthermore, lead peptides with high binding strength and good structural stability were identified from candidate peptides by molecular docking, and molecular dynamics simulation. Conclusions: These results showed that the pipeline with PepPI-DRN, molecular docking and molecular dynamics simulation enabled efficient and reliable identification of lead peptides from the venom gland transcriptome of Orientothele washanensis, providing a robust and effective strategy for the discovery and development of natural peptide drugs from the spider venom.
Xin Zeng, Wen-Feng Du, Wenhao Yin et al.· Pharmaceuticals· 0 citations
The proposed UniPept is a novel model for peptide property prediction that integrates atomic-level and residue-level features to generate robust peptide representations and outperforms competing approaches in interaction and binding affinity prediction tasks.
Hao-Shuang Wu, Yu Deng, Zhen Li et al.· Interdisciplinary Sciences C...· 0 citations
With the worsening crisis of antimicrobial resistance, many researchers are now exploring new types of antibacterial agents, and antimicrobial peptides (AMPs) have begun to attract much attention. However, the experimental identification of AMPs in the large space of natural and synthetic sequences is still slow, expensive and labour-intensive. AMP-Transformer is a two-stage deep learning framework that combines self-supervised pre-training on large-scale unlabelled protein corpora with supervised fine-tuning on curated AMP datasets in this study. The model is a multi-layer bidirectional Transformer encoder that learns contextual residue representations via masked residue modelling, and is then fine-tuned for two coupled tasks: binary discrimination of AMPs from non-AMPs and regression of minimum inhibitory concentration (MIC) values. We used the benchmark datasets built by DBAASP v3, DRAMP 4.0 and other recently published experimental collections for evaluation. On an independent test set, AMP-Transformer had an accuracy of 95.3%, a Matthews correlation coefficient of 0.906, and an area under the receiver operating characteristic curve (AUC-ROC) of 0.986, and outperformed the support vector machine, random forest, convolutional, recurrent and hybrid baselines significantly. Ablation studies show that self-supervised pre-training and multi-head self-attention are the two main contributors, accounting for about 4% of the accuracy increase. Analysis of the learned attention maps shows that the model has independently learned the amphipathic periodicity and cationic residue enrichment characteristic of membrane-active peptides, thereby providing a degree of mechanistic interpretability that is rare among black-box predictors. A subsequent screening of metagenomic open reading frames also produced a ranked list of candidate AMPs with low sequence identity to any training examples, demonstrating the value of the framework for early-stage discovery. Based on the above results, Transformer-based protein language models are relatively stable, interpretable and scalable paradigms for AMP discovery and activity prediction, and they can help establish a practical computational pipeline that selects promising peptide candidates for experimental validation at a lower cost compared with traditional methods.
Shuwen Pan, Eason Soo, Konken Wong· International journal of com...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.