Biopesticidal Potential of the Cannabis sativa L. Metabolites: A Denoised, Docking-Informed QSAR Model
Plant metabolites are a promising source of new biopesticides, but their chemical diversity exceeds the capacity of experimental screening. Cannabis sativa is a particularly attractive crop for discovering such compounds, although its metabolome has not been systematically evaluated for biopesticidal potential. Here, we computationally analyzed 5211 compounds annotated as C. sativa metabolites in the Cannabis Compound Database (CCD) using an integrated framework combining graph-based molecular prediction and protein–ligand interaction analysis. Initial prioritization employed a directed message passing neural network (DMPNN) trained on molecular graphs augmented with RDKit descriptors. The DMPNN predictions were integrated with a CatBoost-derived docking-consistency score based on residue-level Vina interaction terms, reducing the false-positive rate by about 60% compared with the structural model alone. Informative ligand-residue interactions were identified using a random matrix theory (RMT) framework. The DMPNN identified 1010 compounds as DMPNN-positive (score ≥0.70), indicating structural characteristics more consistent with the DS2 pesticide reference set than with the DS3 AChE-inactive reference set. Then, these compounds were filtered using annotations from the CCD to retain 44 secondary metabolites. Finally, the 44 compounds were ranked by the final ensemble score. Compared with reference pesticides, C. sativa metabolites showed higher predicted median oral LD50 values and fewer organ-specific toxicity alerts at the dataset level, although not for all endpoints. Overall, the combined structural and docking-informed workflow identified a small, chemically diverse set of high-ranking C. sativa compounds that can now be prioritized for experimental validation.