By integrating disease-specific context into molecular generation, DrugGen-2 advances AI-assisted drug discovery, offering a powerful tool for de novo design and drug repurposing that accounts for the complex interplay between diseases and molecular targets.
Abstract
Current computational approaches for drug design typically focus on generating molecules conditioned on specific targets or general molecular properties, often neglecting the influence of disease context on target behavior and therapeutic outcomes. To address this gap, we introduce DrugGen-2, a novel generative model that designs small molecules conditioned on both disease ontology and target protein sequences. DrugGen-2 was developed by fine-tuning a pre-trained GPT-2 model on a curated dataset of approved drugs linked to their diseases and targets, using a two-step strategy of supervised fine-tuning followed by reinforcement learning via group relative policy optimization (GRPO). This process was guided by reward functions optimizing for chemical validity, novelty, diversity, and high predicted binding affinity. When evaluated on five protein targets relevant to diabetic nephropathy, DrugGen-2 significantly outperformed baseline models (DrugGPT and DrugGen). It demonstrated a superior capacity to generate unique molecules, exhibited greater structural similarity to approved drugs, and achieved improved predicted binding affinities across all targets. Molecular docking analyses further supported these findings, identifying candidate ligands with strong binding potential, including compounds with predicted affinities (-9.917, -9.485, and -9.367) exceeding those of reference drugs such as enalapril for angiotensin-converting enzyme (-8.283). By integrating disease-specific context into molecular generation, DrugGen-2 advances AI-assisted drug discovery, offering a powerful tool for de novo design and drug repurposing that accounts for the complex interplay between diseases and molecular targets.
This review systematically summarizes the latest developments in heterogeneous graph neural networks, protein language models, and generative artificial intelligence, pointing out the problems currently being addressed in research such as data sparsity and cold start as well as the manifestations of general machine learning challenges.
Qi-Zhong Yang· ITM Web of Conferences· 0 citations
The key innovation of LOGIC lies in constructing a dictionary of functional groups and symptoms, and performing a simple and intuitive multi-hot encoding of drugs and diseases at the micro-scale, and in employing large language models (LLMs) to derive the meso-scale features of diseases without requiring additional domain knowledge.
Yunfei He, Shikai Chen, Yuchen Zhao et al.· IEEE transactions on computa...· 0 citations
Abstract Motivation Identifying drug targets is fundamental in drug development, both for discovering new therapies and for ensuring effective and safe treatments. Drug-target interactions (DTIs) have been predicted using machine learning approaches that integrate heterogeneous data; however, often these models are complex and lack interpretability. Results We investigate whether a simple, fully interpretable linear model can achieve competitive performance for DTI prediction. We propose Linear Interpretable Drug-Target Interaction (LI-DTI), a prediction model inspired by recommender systems. LI-DTI learns from different drug-drug and target-target similarity matrices and provides interpretable predictions as a linear combination of these similarity measures. We show that LI-DTI can recover DTIs even when drugs or targets have no previously known interactions, across multiple cross-validation settings. We further evaluate performance while mitigating potential bias arising from high chemical similarity between drugs or sequence similarity between targets. Finally, we assess LI-DTI in a prospective evaluation, training on DTIs present in DrugBank from 2011 and testing on interactions added through 2022. Across all evaluations, LI-DTI achieves state-of-the-art performance while producing interpretable predictions. For practical use, we provide a web-based tool that enables users to visualize individual LI-DTI predictions for DrugBank (2025) and inspect the biological evidence underlying them. Our results indicate that simple linear models with well-curated similarity features can deliver robust and interpretable DTI predictions, facilitating hypothesis generation and downstream experimental prioritization. Availability and implementation Code and data available at https://github.com/paccanarolab/LI-DTI. Web tool available at https://paccanarolab.org/lidtiweb/.
Santiago Noto, Santiago Ferreyra, Rubén Jiménez et al.· Bioinformatics· 0 citations
GoMA-DTA is proposed, a framework integrating gene ontology (GO) functional annotations with protein semantic features with channelwise gating mechanism that uses functional semantics as anchors to dynamically recalibrate ESM-2embeddings, achieving adaptive semantic filtering.
An Xiong, Zheyu Zhou, Yazi Li et al.· IEEE Transactions on Neural...· 0 citations
BoltzOmics is an interactive, open-source platform that integrates Boltz-2, a deep learning model for biomolecular structure prediction, to rapidly assess mutation effects on drug binding, and establishes a practical AI-driven framework for accelerating computational drug discovery and advancing precision medicine research.
K. Ngo, Kermit L. Carraway, Colleen E. Clancy et al.· iScience· 1 citation
SGTL-DDA is proposed, a novel graph transformer framework designed to incorporate structural information and domain-specific knowledge from heterogeneous biological information networks (HBINs) that successfully identifies both known therapeutics and novel repositioning candidates, supported by molecular docking results and literature evidence.
Bowei Zhao, Hui Zhao, Yu-an Huang et al.· IEEE transactions on computa...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.