Gene expression profiles capture system-level drug responses and offer a promising basis for de novo molecular generation. However, their application is limited by data sparsity and experimental noise, which hinder the reliable mapping between disease-associated transcriptomic perturbations and chemically valid therapeutic molecules. Here, we present AET5, a de novo molecular generation framework that conditions molecular design on disease-reversal gene expression profiles. AET5 integrates contrastive self-supervised learning with pre-trained sequence-to-sequence models to learn robust associations between transcriptomic signatures and molecular structures by deriving noise-tolerant transcriptomic representations and aligning them with molecular sequence space. Across the L1000 dataset, AET5 outperforms existing expression-guided generation methods in generation quality and distributional characteristics, while maintaining favorable physicochemical and drug-related properties. We further apply AET5 to generate candidate compounds for SARS-CoV-2 infection and prostate cancer. Molecular docking and dynamics simulations indicate stable target binding, supporting the biological relevance of the generated molecules. These results demonstrate that disease-reversal expression profiles can effectively guide de novo molecular generation, providing a general framework for biologically informed drug design under noisy transcriptomic conditions.
Zhi-Kang Yuan, Xin Zhang, Gao-Ming Lin et al.· PLoS Computational Biology· 0 citations
INTRODUCTION
Experimental identification of anticancer peptides (ACPs) is timeconsuming and costly, which limits large-scale ACP discovery and screening. To address this challenge, we developed MDFA-MLP, a novel computational framework for ACP prediction that integrates multi-scale feature learning and ensemble classification strategies.
METHODS
The proposed framework combines physicochemical descriptors, including amino acid composition (AAC), dipeptide composition (DPC), composition-transition-distribution (CTD), and pseudo-amino acid composition (PseAAC), with ProtBERT-derived embeddings. A Multiscale Dilated Fusion Attention (MDFA) module was designed to capture sequence patterns at different scales and enhance feature fusion. An ensemble classifier consisting of a multilayer perceptron (MLP), support vector machine (SVM), and histogram-based gradient boosting (HGB) was employed to improve prediction robustness and stability.
RESULTS
The proposed model was evaluated on the AntiCP 2.0 dataset under the different negative-sample settings. On Dataset A, MDFA-MLP achieved an accuracy of 93.9%, sensitivity of 92.2%, specificity of 95.8%, and an MCC of 0.89. On the more challenging Dataset B, the model achieved an accuracy of 78.2%, sensitivity of 76.5%, specificity of 82.6%, and an MCC of 0.65. Comparative experiments demonstrated that MDFA-MLP achieved competitive and balanced performance across multiple evaluation metrics. Although the improvement over existing methods was moderate in some cases, the model maintained stable predictive performance under different negative-sample settings, indicating good robustness and generalization ability.
DISCUSSION
The results indicate that traditional sequence descriptors and deep protein language model embeddings provide complementary biological information. The MDFA module effectively enhances feature representation by integrating multi-scale sequence characteristics, while the ensemble strategy improves model robustness and generalization under varying data distributions.
CONCLUSION
MDFA-MLP provides an effective and reliable framework for ACP prediction. By integrating handcrafted descriptors, protein language model representations, and ensemble learning, the proposed method can facilitate large-scale computational screening of candidate ACPs prior to experimental validation.