Skip to content

Author

Dr. Pradyumna Kumar Pradhan

1 paper indexed here

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

#graph neural networks Open access Sep 2026

Plasma Protein Binding Prediction under Scaffold Split: Tree Models, ChemBERTa, Graph Networks, and Late Fusion

From idea generation to testing, drug discovery is a lengthy and expensive process. Rather than chasing all of their molecule ideas, scientists prioritize certain molecules that are more likely to be successful, from synthesis to safety to therapeutic use. Machine learning and AI enable researchers to narrow their candidate choices by predicting important properties such as ADMET (Absorption, Distribution, Metabolism, Excretion, and Toxicity). This study uses the Plasma Protein Binding (PPB) dataset from DeepChem’s MoleculeNet, with several graph neural networks (GNNs), ChemBERTa (a language model for SMILES strings of molecules), and fusions of these models. The performance of these neural models is compared against two standard tree-based models: Random Forest (RF) and XGBoost. Instead of random splitting of the dataset, this study evaluates the models based on scaffold splitting that separates molecules by the Bemis–Murcko scaffold and puts the entire group into only the training, validation, or test set. This study demonstrates that for the PPB dataset, GNNs and ChemBERTa models perform better than the tree models. On a mean basis, the fusion models perform better than the individual GNNs or ChemBERTa-only models. Random splitting on the RF model exhibits a better R2 score than on a scaffold-split dataset. Because under scaffold splitting the model predicts on unseen cores, the task becomes more difficult than under regular random splitting.

Dr. Pradyumna Kumar Pradhan · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.