Skip to content

Benchmarking Experiments with External Data, Domain-Specific Features, and Learning Strategies for Modeling ADME and Potency of Coronavirus Main Protease Inhibitors

Sep 2026 · Journal of Chemical Information and Modeling · 0 citations · 86 references
Computational Drug Discovery Methods

TL;DR

This work tested the key molecular descriptors capable of identifying antivirals, augmenting challenge data sets with curated public data, and building both classical machine learning and graph neural network (GNN) variants, providing both practical benchmarked modeling strategies for ADME and potency prediction and a reproducible computational foundation for AI-driven antiviral drug discovery efforts.

Abstract

Predictive computational models of potency and pharmacokinetic properties of antiviral drug candidates can dramatically accelerate the path from molecular design to clinical evaluation. The ASAP-Polaris-OpenADMET consortium hosted the open-science Antiviral Drug Discovery Challenge 2025, providing a rigorous setting for this problem and requiring prediction of potency against the SARS-CoV-2 and MERS-CoV Mpro alongside five ADME end points (LogD, KSOL, HLM, MLM, and MDR1-MDCKII). In response, we pursued a domain-aware modeling strategy, reasoning that the distinctive chemical properties of antiviral compounds could be explicitly engineered in the feature space. We tested this premise by isolating the key molecular descriptors capable of identifying antivirals, augmenting challenge data sets with curated public data, and building both classical machine learning and graph neural network (GNN) variants. For ADME modeling, ensemble gradient boosting models trained on our bespoke feature space consistently outperformed both broader tree-based approaches and GNN variants across the five ADME end points in blind evaluation. By a factorial design of experiments and statistical benchmarking, we evaluated the performance of the optimal XGBoost model for each end point. For modeling potency, we explored varied pretraining paradigms and ensemble fusion strategies, and found that a stacking ensemble combining docking-enhanced XGBoost with a stereochemistry-aware AttentiveFP Graph-Transformer model yielded the strongest significant generalization on the Challenge test set, suggesting that orthogonal model families reduced epistemic uncertainty attributable to model selection in data-limited, multitarget settings. Beyond benchmarking, our experiments provided a testbed for evaluating critical assumptions in the field, encompassing data augmentation strategies, splitting protocols, and ensemble integration approaches. Our analysis revealed that a consistently optimal model is elusive, and that several standard industry practices fail to withstand close inspection. Overall, our work provides both practical benchmarked modeling strategies for ADME and potency prediction and a reproducible computational foundation for AI-driven antiviral drug discovery efforts.

View source

Similar papers

#computer vision Conference Aug 2008

Scrum in a Multiproject Environment: An Ethnographically-Inspired Case Study on the Adoption Challenges

Agile methods continue to gain popularity. In particular, the Scrum method appears to be on the verge of becoming a de-facto standard in the industry, leading the so called Agile movement. While there are success stories and recommendations, there is little scientifically valid evidence of the challenges in the adoptio...

A. Marchenko, P. Abrahamsson · 59 citations · ⚡11
#computer vision Open access Sep 2012

Making the leap to a software platform strategy: Issues and challenges

A comprehensive taxonomy of the challenges faced when a medium-scale organization decided to adopt software platforms is provided, namely: business challenges, organizational challenges, technical challenges, and people challenges.

Yaser Ghanam, F. Maurer, P. Abrahamsson · 41 citations · ⚡3
#machine learning Open access Mar 2024

Integration of molecular coarse-grained model into geometric representation learning framework for protein-protein complex property prediction

MCGLPPI, a novel geometric representation learning framework that combines graph neural networks (GNNs) with the MARTINI molecular coarse-grained (CG) model to predict overall PPI properties accurately and efficiently, offers an effective and efficient solution for PPI overall property predictions.

Yang Yue, Shu Li, Yihua Cheng et al. · 15 citations

PepPCBench is a Comprehensive Benchmarking Framework for Protein-Peptide Complex Structure Prediction

PepPCBench enables a robust evaluation of PFNN-based methods and supports their continued development for peptide-protein structure prediction, and highlights the influence of peptide length, conformational flexibility, and training set similarity on prediction accuracy.

Si-Long Zhai, Huifeng Zhao, Ji-Ke Wang et al. · 13 citations · ⚡1
#machine learning Open access Sep 2025

Unified and explainable molecular representation learning for imperfectly annotated data from the hypergraph view

OmniMol is presented, a framework using hypergraphs to improve predictions of molecular properties, addressing challenges of imperfect data annotation and enhancing model explainability, and achieves state-of-the-art performance in properties prediction.

Bowen Wang, Junyou Li, Donghao Zhou et al. · 11 citations

Related blog posts

Microsoft Research Blog Jul 13, 2026

Verifying Rust cryptography in SymCrypt, from standards to code

Cryptographic code supports vital protections in modern computing systems. Learn how a new method helps verify code as developers write it while preserving speed and adaptability as it gets implemented and evolves. The post Verifying Rust cryptography in SymCrypt, from standards to code appeared first on Microsoft Research.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.