Aug 2026· Combinatorial chemistry & high throughput screening· 0 citations
Medicine
Abstract
INTRODUCTION
Experimental identification of anticancer peptides (ACPs) is timeconsuming and costly, which limits large-scale ACP discovery and screening. To address this challenge, we developed MDFA-MLP, a novel computational framework for ACP prediction that integrates multi-scale feature learning and ensemble classification strategies.
Methods
The proposed framework combines physicochemical descriptors, including amino acid composition (AAC), dipeptide composition (DPC), composition-transition-distribution (CTD), and pseudo-amino acid composition (PseAAC), with ProtBERT-derived embeddings. A Multiscale Dilated Fusion Attention (MDFA) module was designed to capture sequence patterns at different scales and enhance feature fusion. An ensemble classifier consisting of a multilayer perceptron (MLP), support vector machine (SVM), and histogram-based gradient boosting (HGB) was employed to improve prediction robustness and stability.
Results
The proposed model was evaluated on the AntiCP 2.0 dataset under the different negative-sample settings. On Dataset A, MDFA-MLP achieved an accuracy of 93.9%, sensitivity of 92.2%, specificity of 95.8%, and an MCC of 0.89. On the more challenging Dataset B, the model achieved an accuracy of 78.2%, sensitivity of 76.5%, specificity of 82.6%, and an MCC of 0.65. Comparative experiments demonstrated that MDFA-MLP achieved competitive and balanced performance across multiple evaluation metrics. Although the improvement over existing methods was moderate in some cases, the model maintained stable predictive performance under different negative-sample settings, indicating good robustness and generalization ability.
Discussion
The results indicate that traditional sequence descriptors and deep protein language model embeddings provide complementary biological information. The MDFA module effectively enhances feature representation by integrating multi-scale sequence characteristics, while the ensemble strategy improves model robustness and generalization under varying data distributions.
Conclusion
MDFA-MLP provides an effective and reliable framework for ACP prediction. By integrating handcrafted descriptors, protein language model representations, and ensemble learning, the proposed method can facilitate large-scale computational screening of candidate ACPs prior to experimental validation.
Accurate computational prediction of antiviral peptides (AVPs) can accelerate peptide screening and reduce experimental costs. However, existing deep learning-based methods still suffer from severe class imbalance, over-reliance on handcrafted features and limited interpretability. Here, we propose PMAVP, a multi-task learning framework that integrates the ProtT5 pre-trained protein language model with a Mamba-inspired module for AVP identification and functional activity prediction. We use ProtT5 to extract deep semantic representations from peptide sequences and a Mamba module to capture long-range dependencies at a lower computational complexity. We introduce Focal Loss to mitigate class imbalance and leverage transfer learning to enhance performance on functional activity prediction. Experimental results demonstrate that our model achieves superior performance in terms of prediction accuracy, stability, and computational efficiency. Furthermore, DeepSHAP-based interpretability analysis reveals that the first 40 amino acid residues contribute substantially to AVP prediction.
Pei-Wei Wei, Wei-Hao Su, Qing-Song Qin et al.· International Journal of Mol...· 0 citations
Experimental results on multiple benchmark data sets demonstrate that MMU-DPI outperforms several state-of-the-art DPI prediction methods and indicate that MMU-DPI can serve as a useful computational tool for drug discovery.
Jiahao Wei, Tie Shen· Journal of Chemical Informat...· 0 citations
The results show that the presented MSFGCN model is very accurate and can be used reliably for ligand docking prediction and proves useful in future for rapid drug discovery and computational drug design.
P. Shunmugaraj, ·. T. F. Abbs, Fen Reji· Journal of the Iranian Chemi...· 0 citations
MOTIVATION
As unique drugs positioned between small and macro molecules, anticancer peptides (ACPs) hold great potential in oncotherapy owing to their high selectivity and low toxicity. Nowadays, computational ACP prediction has emerged as a cost-effective alternative to bioassay screening, but most methods are limited to identifying bioactivity and fail to resolve tumor cell-specific targeting, primarily because of the sparse annotated data.
RESULTS
To fill this gap, we integrate a hybrid dataset compiled from five well-established peptide databases and propose TargetPC, a deep learning method tailored for cell line-targeted ACP prediction. TargetPC encodes multimodal representations of ACPs and cell lines via pretrained protein and omics models, and combines them via hierarchical intra- and inter-modal fusion for targeting prediction. This combination is further augmented by a mutual learning paradigm that distills domain knowledge from both ACP and cell line, enabling improved generalization under sparse supervision. Experimental results on the hybrid dataset demonstrate the effectiveness of TargetPC, which outperforms the state-of-the-art baselines in terms of prediction accuracy, and maintains strong generalization to unseen ACPs and cell lines. When extended to out-of-distribution samples, TargetPC has successfully screened dozens of novel ACPs targeted to breast cancer cells and uncovered biological motifs underlying its predictions. As a result, our TargetPC is expected to serve as a versatile tool for lead ACP discovery at a lower burden.
AVAILABILITY
The source code and data are available at GitHub (https://github.com/liuxuan666/TargetPC).
Xuan Liu, Jian Zhang, Chong-Yang Chen et al.· Bioinformatics· 0 citations
Predicting peptide retention time (RT) remains a significant challenge, particularly when training data is limited. In this study, we present MetaRT, a stacked-ensemble machine learning framework designed to predict the RTs from small dataset of hydrophobic peptides. Peptides composed of hydrophobic amino acids—phenylalanine (F), isoleucine (I), methionine (M), and tryptophan (W) were synthesized, and their experimental RTs were measured from the mixture entities. MetaRT utilizes a graph convolutional network (GCN) to extract structural features from the peptide sequences. The MetaRT model architecture employed multiple base learners, integrating the outputs through a meta-learner optimized via hyperparameter tuning and 3-fold cross-validation. Besides, the performance of MetaRT was compared to three ensemble methods - weight averaging, bagging, and boosting. The results demonstrated that structure-based MetaRT outperformed both base learners and the ensemble models, achieving a lower root mean square error (RMSE) of 0.08 and a maximum RT deviation of approximately 1.4 min. Compared to the prediction performance on molecular descriptors inclusion, the structure-guided model consistently performed well in terms of RMSE. Notably, MetaRT accurately predicted the RTs of sequence isomers by leveraging the structural features, with deviations ranging from 0.2 to 1.3 min. In contrast, descriptor-based model showed increased prediction error for the isomeric sequences. For peptides with lower hydrophobicity that were not included in the training data, the structure-based predictions led to the maximum deviation of 4.9 min from the experimental RTs. The entire predicted RTs were subsequently validated by linear regression analyses with the corresponding experimental values. These findings highlight the potential of MetaRT as a structure-based predictive tool for improving RT prediction accuracy, especially in data-limited scenarios. Future work will focus on enhancing the robustness of MetaRT by incorporating a wider variety of peptide classes to further refine its predictive capabilities.
R. A. Mahmood, M. Mahin, R. Alam et al.· Journal of Analytical Scienc...· 0 citations
Overall, DrugPLMFormer provides a reproducible, leakage-aware framework for retrospective sequence-based druggability screening and target prioritization, while prospective validation and experimental confirmation remain necessary before operational deployment.
Z. Kafi, Khosro Rezaee, Hossein Eslami· Journal of King Saud Univers...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.