The central challenge in de novo protein design is generating plausible, mutually compatible structures and sequences, such that each designed sequence folds into its intended structure and the structure accommodates that sequence. Compared to typical two-stage design methods, which decouple the modeling of the interde...
Yuan-Le Mo, Bo Qiang, Hai-Tao Lin et al.· 0 citations
RFO formulates binder improvement as a residue-wise mutational search problem, sampling candidate substitutions alternately based on gradient-guided sequence optimization using all-atom structure prediction models and a cycling-based sequence redesign strategy that alternates structure generation with an orthogonal pre...
Odin Zhang, Jia-Qi Wang, T. Thompson et al.· bioRxiv· 0 citations
This study presents ECloudGen, which uses latent diffusion to generate electron clouds from protein pockets and decodes them into molecules, and adopts two-stage training, which expands the chemical space accessible to generative drug design.
RAPiDock is presented, an all-atom diffusion model that predicts peptide–protein binding patterns across 92 amino acid types, enabling high-throughput virtual screening for advancing therapeutic peptide design and serve as a powerful tool for high-throughput virtual screening with structural precision.
Accurate prediction of metalloprotein-ligand interactions is critical for metalloprotein-targeted drug discovery. Conventional docking tools and existing deep learning (DL) models fail to reliably capture metal-ligand interactions, hampering the discovery of potent metalloprotein inhibitors. Here, we propose MetalloDoc...
Hui Zhang, Xujun Zhang, Qun Su et al.· Journal of the American Chem...· 6 citations
The results indicate that this fully automated, open-source system holds potential value for improving the efficiency and sustainability of molecular synthesis, and the integration of organic and enzymatic synthesis enhances molecule construction efficiency.
Token-Mol is presented, a token-only 3D drug design model that encodes both 2D and 3D structural information, along with molecular properties, into discrete tokens, which introduces a Gaussian cross-entropy loss function tailored for regression tasks, enabling superior performance across multiple downstream application...
Ji-Ke Wang, Rui Qin, Mingyang Wang et al.· Nature Communications· 30 citations· ⚡1
ERAM aligns pre-trained molecular representations from Protein Language Model with the knowledge of enzyme catalysis by modeling enzymatic reactions as multi-relational data, and demonstrates its potential as a versatile and effective tool for enzyme catalysis research.
The results show that current LLMs capture partial epitope-related signals but remain limited in antibody-specific sequence grounding, long-context residue localization, and biologically grounded reasoning, so EpiBench provides a diagnostic testbed for measuring and improving sequence-aware biomedical LLMs toward relia...
Zi-Rui Wang, Jiaqing Wang, Qing-Han Wang et al.· 0 citations
ProphDR is an interpretable deep learning framework that integrates multiomics data and drug structural information using a hierarchical attention mechanism, and generates biologically interpretable attention maps that highlight key pharmacophores and resistance-related genes consistent with established mechanisms in N...
Yundian Zeng, Qing Ye, Jike Wang et al.· Journal of Chemical Informat...· 0 citations
The Comprehensive VS Platform with AI Engine (CVSP-AIE) for drug discovery from compound libraries integrates three AI models: KarmaDock, a fast docking model that directly updates atomic coordinates; CarsiDock, an accurate docking model that predicts protein-ligand distances and reconstructs binding poses; and RTMScor...
NACraft, a training-free and programmatic framework for all-atom nucleic-acid aptamer design based on backpropagation through structure-model feedback, is presented, demonstrating the effectiveness and versatility of NACraft and extending structure-model hallucination toward programmatic nucleic-acid aptamer design.