Jul 2026· Chemical Research in Toxicology· Vol 39, pp. 1518 - 1531· 0 citations· 50 references
Medicine
TL;DR
A dedicated Python-based computational package designed for the systematic development, training, and evaluation of ML models with an explicit consideration of interspecies data integration in ADMET modeling, offering quantitative guidance on when and how cross-species and cross-assay data improve predictive performance.
Abstract
The rapid expansion of in silico methodologies has reshaped modern drug discovery and toxicology research; however, robust prediction of ADMET endpoints remains limited by the scarcity and heterogeneity of experimental data. In particular, toxicological data sets are often fragmented across species, complicating the development of reliable and generalizable machine learning models. To address this challenge, we introduce a dedicated Python-based computational package, ADMET-XSpec, designed for the systematic development, training, and evaluation of ML models with an explicit consideration of interspecies data integration. The framework enables controlled incorporation of chemical space originating from different species and different assay types, allowing users to flexibly construct single-species models, as well as models augmented with cross-species information. This design facilitates systematic investigation of how additional data from other organisms influence model performance without imposing assumptions inherent to specific transfer learning paradigms. By supporting standardized preprocessing, scalable integration of heterogeneous data sets, and rigorous benchmarking, the proposed tool provides a unified environment for studying cross-species effects in ADMET modeling. Overall, this work delivers a practical resource for the ADMET modeling community and offers insights into how interspecies and interassay data integration can improve model robustness and generalizability while clarifying the conditions under which cross-species and cross-assay data information is beneficial for predictive toxicology. The package is freely available at https://github.com/hubertrybka/admet-xspec. ADMET-XSpec advances the state of the art by providing the first dedicated framework for controlled interspecies and interassay data integration in ADMET modeling, offering quantitative guidance on when and how cross-species and cross-assay data improve predictive performance.
Abstract Motivation Chemical toxicity assessment is critical for drug development and environmental safety. Computational models have emerged as a promising alternative to animal testing and now play a significant role in efficiently evaluating new chemicals. To address the urgent need for user-friendly machine learning tools in computational toxicology, we developed ToxiVerse, a public web-based platform. Results ToxiVerse provides automatic chemical bioprofiling, curated toxicity datasets, and a predictive modeling interface designed for researchers who lack programming expertise. The platform comprises three integrated modules: (i) Bioprofiler, which provides chemical descriptors by combining chemical-bioactivity data from PubChem assays with a machine learning-based data gap-filling procedure; (ii) Database, which hosts ∼50 000 curated chemicals covering diverse toxicity endpoints; and (iii) Cheminformatics, which enables dataset upload, chemical curation, and automatic generation of quantitative structure–activity relationship models for toxicity prediction. Availability The tool is accessible at www.toxiverse.com, and source code is available at https://github.com/zhu-research-group/toxiverse.
Prasannavenkatesh Durai, Daniel P. Russo, Yitao Shen et al.· Bioinformatics· 0 citations
Drug discovery is frequently limited by high attrition rates, and poor absorption, distribution, metabolism, excretion, and toxicity (ADMET) profiles are a major cause of late-stage failure. Therefore, precise ADMET property prediction is necessary to develop safe and effective drug candidates. Traditional experimental assays and rule-based computational procedures are limited by their poor predictive power, cost, and time, despite providing valuable insights. Innovative strategies to deal with these issues have been introduced by developments in artificial intelligence (AI), such as machine learning (ML), deep learning (DL), graph neural networks (GNNs), generative models, and multi-task learning (MTL). AI techniques can better generalize scaffolds, capture interdependencies between pharmacokinetic and toxicological endpoints, and model complex nonlinear relationships by leveraging large, diverse datasets. Explainable AI (XAI) enhances transparency by detecting biological and structural characteristics that are relevant to predictions, even if integrated pipelines combine predictive modeling with molecular creation and optimization. AI-driven ADMET prediction is becoming a vital tool in lowering attrition, speeding up candidate prioritization, and influencing the direction of rational drug development, despite persistent issues with data quality, regulatory acceptance, and synthetic viability.
Satyam Kumar Vishwash, Ram Babu Soni, Ratima Sood et al.· Current Computer - Aided Dru...· 0 citations
This work presents a model validation framework consisting of five recommendations that would enable the community to move beyond aggregate metrics toward understanding where and why molecular property prediction models fail, and connects evaluation choices to real-world applications and case studies encountered in pharmaceutical research.
Srijit Seal, Akshat Shirish Zalte, David Alencar Araripe et al.· bioRxiv· 0 citations
This tutorial provides a comprehensive, end-to-end workflow from raw data to deployed models,icitly designed for environmental chemists with limited prior experience in ML modeling while also providing practical guidance for other users seeking to strengthen their modeling workflows.
Kai Zhang, Yushu Cheng, Hai-Ping Ai et al.· ACS Environmental Au· 0 citations
Monroe is presented, a new MFM with several innovations over the existing state of the art: increased scale allowing pre-training on over 81 million molecules from the PM6 quantum chemistry dataset; improved graph representation of stereochemistry; improved training losses including conformer denoising and embedding decorrelation; improved multi-task learning; and the use of a prior-data-fitted model (TabPFN) for downstream in-context prediction.
Blazej Banaszewski, Andrew W. Fitzgibbon· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.