Skip to content
Conference

AI-Driven Virtual Screening of Naturalcompounds Against SARS-CoV-2 Using Embedding-Based Drug-Protein Interaction Prediction

Jul 2026 · 2026 6th International Conference on Electrical, Computer and Energy Technologies (ICECET) · pp. 1-6 · 0 citations · 21 references

Abstract

The imperative necessity for rapid discovery of antiviral agents against emerging viral diseases, such as COVID-19 caused by SARS-CoV-2, has emphasised the limitations of conventional drug discovery, with its deliberate pace, high costs, and high failure rate. In this study, we present a high-throughput “AI-driven virtual screening pipeline” that combines cutting-edge molecular embeddings using transformers for compounds and language model-based protein sequences for targets, and combines these using a gradient-boosted decision tree-based model (XGBoost) trained on ChEMBL bioactivity data, with excellent performance for binary prediction (ROC-AUC 0.8408, Accuracy 0.76) comparable to top-performing methods. When applied to the large library of natural products from COCONUT database, our model correctly predicted top-ranked compounds, some of which were shortlisted for molecular docking against the SARS-CoV-2 spike receptor-binding domain (RBD), the key interface for ACE2 binding, in order to test the validity of proposed pipeline and the results revealed several promising compounds with excellent binding energies and multiple interactions with hotspot residues such as K417, Y453, Q493, G496, Q498, N501, Y505 and F486. The potential of cost-effective, synergistic integration of scalable deep learning of representations, interpretable machine learning, and physics-based refinement of structure is significant for accelerating natural product-based therapeutics for coronaviruses and possible other viral threats, with the potential to expand to ensemble approaches, active learning, and variant-based targets for enhanced efficacy.

View source

Similar papers

Open access Jul 2026

Deep learning-based prediction of drug-target interactions between antiviral drugs and SARS-CoV-2 proteins using an image-based representation approach.

This study applies MPS2IT-DTI (Molecule and Protein Sequence to Image Transformer for Drug-Target Interaction), a deep learning framework that represents molecular and protein sequences as images as images using k-mer frequency encoding, enabling convolutional neural networks to capture spatial compositional patterns associated with biochemical interactions.

Jackson G. de Souza, Marcelo A. C. Fernandes, Raquel de Melo Barbosa · 0 citations
Open access Aug 2026

Discovery of two novel small-molecule series with potent SARS-CoV-2 inhibition via putative nsp13 open-state trapping

We report the discovery of 31 SARS-CoV-2 inhibitors identified across two computational-to-experimental screening campaigns, with average cell-based antiviral potency satisfying 〈EC50〉 ≤ CC50/3 in A549-hACE2 cells. These compounds were selected from 60 successfully tested molecules evaluated using nanoluciferase reporter assay (nLuc), cytopathic effect (CPE), and host-cell cytotoxicity (CC50) assays. We used a funnel-like computational framework where candidate compounds were progressively prioritized through different scores, improving robustness against the limitations of individual metrics. However, none of the docking-based scoring metrics, including MM/GBSA and Glide scores, showed statistically significant correlation with experimental activity, illustrating the limitations of docking-only approaches and supporting a major contribution of the machine-learning predictions to the high hit-identification rate. Although direct target validation is still lacking, the machine-learning (ML) workflow was grounded in experimental nsp13 inhibition labels, and molecular dynamics (MD) simulations suggest that these compounds can act by stabilizing apo-like open conformations of nsp13, with differential interactions involving the 1B domain contributing to potency differences. Among the compounds evaluated, the phenoxypropanol (PP) and bipiperidine (BPP) series stood out as the most promising SARS-CoV-2 antivirals. Integrated analysis of structure–activity relationships, physicochemical parameters, metabolic stability, and interdomain dynamics delineated structural optimization paths for both series.

Alma C. Castañeda-Leautaud, Thomas D. Bannister, Eunjung Kim et al. · 0 citations
Jul 2026

Real-World Assessment of Machine-Learned Docking Using Bioassay-Derived Benchmarks

This work systematically evaluates the performance of a popular ML-based docking method, DiffDock-Pocket, on high-throughput screening (HTS) data sets derived from the PubChem BioAssay database, a premier source of bioactivity data.

Furyal Ahmed, M. Soellner, Charles L. Brooks · 0 citations
Open access Sep 2026

AI-enhanced adaptive virtual screening of large libraries for ligand discovery.

Ultralarge virtual screenings (ULVSs) evaluate billions of molecules for drug discovery but face cost, flexibility and scalability limits. We introduce AdaptiveFlow, an open-source platform that makes ULVSs more accessible, scalable and efficient and supports artificial intelligence (AI) and machine learning (ML) method development. AdaptiveFlow provides a screening-ready version of the Enamine REAL Space, to our knowledge the largest library of ready-to-dock, drug-like molecules, comprising 69 billion compounds, also available in SELFIES format. An 18-dimensional grid of molecular properties prioritizes promising chemical subspaces, with optional active learning, reducing computational costs by orders of magnitude. AdaptiveFlow integrates >1,500 docking protocols, including GPU-accelerated and ML-based methods, and achieves near-linear scaling on up to 5.6 million CPUs in the Amazon Web Services cloud. We identified nanomolar inhibitors of two disease-relevant targets, ferroptosis suppressor protein 1 (FSP1) and poly(ADP-ribose) polymerase 1. Co-crystal structures provided mechanistic insights into FSP1 inhibition. AdaptiveFlow enables drug discovery at unprecedented scale and supports the development of AI-driven methods.

Domiziana Cecchini, AkshatKumar Nigam, Ming Tang et al. · 0 citations
#graph neural networks Review Open access Aug 2026

Bridging the antiviral drug design gap: a combined machine learning and QSAR approach for drug repurposing of host kinase inhibitors

Viral outbreaks combined with rapid emergence of mutated viruses have highlighted an urging need for accelerating antiviral drug discovery pipelines. Unfortunately, current drug discovery remains stuck to conventional methods which are slow especially during pandemics. In this article, we present a literature-based synthesis of an integrative machine learning (ML) guided QSAR framework that unifies ligand-based, structure-based, and systems biology approaches towards the aim of generating a host directed antiviral repurposing strategy. Moreover, a modern ML enhanced QSAR modeling strategy is proposed to target host directed therapeutics (HDTs), particularly the host kinase enzymes. The proposed framework integrates molecular descriptor modeling, ensemble learning methods (e.g., RF, gradient boosting), graph neural networks (GNNs), and multi-omics target prioritization to outline a predictive antiviral repurposing model. This structured workflow encompasses dataset assembly, descriptor generation, model training, virtual screening, and experimental validation as sequential stages to guide, rather than as a pipeline that has itself been built or independently validated here, translational deployment. The review is illustrated through a retrospective narrative synthesis of four independently published, clinically relevant repurposed HDTs, namely Baricitinib, Lapatinib, Bemcentinib, and Sunitinib. These published case studies, drawn from the primary literature, exemplify how AI/ML-enhanced QSAR and network-based approaches have been used elsewhere to identify active antiviral kinase inhibitors; they are presented here as illustrative evidence of feasibility of such a computational pipeline. Thus, the AI guided repurposing of host kinase inhibitors offers a systematically accelerated strategy to bridge the drug design gap, with the potential for faster therapeutic deployment against viral threats pending prospective, harmonized validation. This review describes a framework that combines artificial intelligence (AI), machine learning (ML), and Quantitative Structure-Activity Relationship (QSAR) modeling to speed up the search for new antiviral drugs. Instead of targeting the virus directly, the framework targets host cell proteins such as kinases, which many viruses hijack during infection, an approach also known as host-directed therapy (HDT). To show how this approach could work, we review four drugs that were originally developed for other diseases and later found to also fight viral infections: Baricitinib, Lapatinib, Bemcentinib, and Sunitinib. Each case was reported independently in the published literature, and we present them here as examples of what AI-assisted drug repurposing can achieve, not as proof that our specific framework has itself been built and tested. Accelerated therapeutic antiviral drug discovery pipelines are being a critical need due to viral outbreaks and rapid emergence of mutated viruses. Host Directed Therapeutics (HDTs) are new drug discovery strategies that can modulate specific host pathways essential for viral multiplication. The AI-HDT Framework is proposed to bridge the gap between the computational chemical prediction and clinical real-life application. The integration of multi-omics data such as phosphoproteomic data and transcriptomic data using the GNN models will help scientist to identify uniquely expressed host genes during the various episodes of viral infection.

Rand O. Shahin, Yusra Azzam, Salma Azzam · 0 citations
Open access Jul 2026

Machine Learning Integrated Designing and Screening of 8-Hydroxyquinoline Based Metallo-β-Lactamase Inhibitors

The rapid emergence of metallo-b-lactamase-mediated antibiotic resistance has created an urgent need for new inhibitor discovery strategies. In this work, a machine-learning-guided workflow was developed to generate and prioritize potential inhibitors targeting NDM-1. A SMILES-based variational autoencoder was first pretrained on a broad molecular dataset to learn general chemical syntax and latent molecular representations. The model was then fine-tuned on an 8-hydroxyquinoline-enriched dataset to bias molecular generation toward zinc-binding chemical space relevant to metallo-β-lactamase inhibition. Generated compounds were processed through structural filtering and docking-based evaluation to create training data for downstream predictive modeling. Molecular fingerprints and physicochemical descriptors were then used to train XGBoost models for docking score prediction and classification of potential binders. Classification proved especially useful for prescreening because it avoided overinterpreting small differences in noisy docking scores while still enriching for compounds likely to perform well in docking. The resulting workflow demonstrates how generative modeling and supervised machine learning can be combined to reduce chemical search space, prioritize candidate inhibitors, and guide computational drug discovery. Although experimental validation remains necessary, this approach provides a scalable framework for identifying promising zinc-binding compounds for further molecular simulation and inhibitor development that can be expanded in future studies.

Anthony M. Baudino, Kari L. Stone · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.