Jun 2026· Journal of Chemical Information and Modeling· 1 citation· 63 references
Medicine
TL;DR
This study systematically compares molecular representations (Mordred, ECFP, ChemBERTa, UniMol2, etc.) across several machine learning architectures to predict CD-PFAS binding energies, demonstrating that molecular representation choice is critical for small-data host-guest binding prediction.
Abstract
Per- and polyfluoroalkyl substances (PFAS) persist in water systems and resist conventional removal methods such as activated carbon, which shows reduced efficiency with short-chain PFAS and in the presence of dissolved organic matter. Cyclodextrin-based polymers (CDPs) have emerged as sustainable alternatives, with competitive and selective PFAS adsorption capabilities. These polymers consist of glucose-based cyclodextrin (CD) units that can form host-guest inclusion complexes with PFAS pollutants. However, these binding interactions are not fully understood or quantified. We conducted an evaluation of machine learning approaches to model these host-guest interactions, providing insights into predictive capabilities for later CDP design. This study systematically compares molecular representations (Mordred, ECFP, ChemBERTa, UniMol2, etc.) across several machine learning architectures to predict CD-PFAS binding energies. First, we generated molecular embeddings of 3459 experimental host-guest pairs in the OpenCycloDB data set and 63 external CD-PFAS pairs. We then compared these embeddings via AlignedUMAP visualizations and nearest neighbor analyses. Next, we trained and evaluated predictive models using these embeddings on the OpenCycloDB data set, exploring the effectiveness of transfer learning and finetuning techniques. We finally tested model generalizability on two external experimental CD-PFAS binding data sets. All embeddings captured relevant chemical features, where UniMol2 differed most from other methods in embedding space analysis. Predictive models performed variably based on embedding choice and architecture, with the best-performing combination achieving moderate accuracy on the OpenCycloDB data set. Embeddings pretrained on large molecular data sets and finetuning the ChemBERTa embeddings both showed predictive improvements. However, external validation revealed limited generalizability to CD-PFAS complexes, highlighting domain shift challenges. Notably, leave-one-out cross-validation on the external PFAS-specific data indicated that training on in-domain data improved predictive performance at the cost of generalizability. This work demonstrates that molecular representation choice is critical for small-data host-guest binding prediction. However, domain shift between general CD data and specialized CD-PFAS applications remains a fundamental challenge, for which transfer learning and finetuning may offer potential solutions for future data-driven pipelines for CDP design and sustainable PFAS removal.
The enzymatic degradation of poly(ethylene terephthalate) (PET) offers a sustainable route for plastic recycling but is often hindered by limited enzyme adsorption on hydrophobic surfaces. Inspired by carbohydrate-binding modules (CBMs), which enhance enzyme performance on insoluble substrates, we developed a machine-learning-assisted pipeline to discover PET-binding modules from natural protein architectures. Integration of profile hidden Markov model-based homology searching with a supervised PET hydrolase machine-learning model (PETML) revealed that CBMs belonging to family 13, typically known for glycan recognition, were the most abundant CBMs associated with putative PET hydrolase homologs in the screened dataset. From 197 non-redundant candidates, high-throughput docking and molecular dynamics simulations prioritized tCBM13-1 (WP_357125140.1), which exhibited stable interfacial binding via cooperative aromatic and polar interactions. When fused to sfGFP, tCBM13-1 demonstrated superior adsorption (∼70%) and surface retention (∼90%) on PET powder at 37 °C and 45 °C, outperforming a benchmark CBM2. Co-displayed with FAST-PETase on the Escherichia coli (E. coli) surface using a dual-anchor system (OmpA and EhaA), tCBM13-1 enhanced PET film depolymerization by ∼ 43%, achieving a rate of 2793 μg/(d·cm2). The whole-cell catalyst retained > 64% of its initial activity after 10 cycles, indicating robust recyclability. This work integrates machine-learning-guided module mining with synthetic biology to engineer efficient, reusable biocatalysts for PET degradation, offering a scalable strategy for polymer biorecycling.
Rui Long, Yaxin Tang, Chengyong Wang et al.· Bioresource Technology· 1 citation
Metal–organic frameworks (MOFs) are versatile materials with tunable crystal structures, morphologies, and chemistries, offering diverse physical and chemical properties. Although typically electrically insulating, specific combinations of organic and inorganic components can impart electrical conductivity to MOFs. The virtually limitless chemical space of MOFs, however, presents a significant challenge in identifying optimal candidates for various applications. Although density functional theory (DFT) can probe their electronic structure, its high computational cost hinders the discovery of novel electroactive MOFs using machine learning due to limited data. To tackle these challenges, a semiempirical extended tight-binding approach (GFN1-xTB) is employed to compute the electronic properties of a dataset of MOFs, and it is shown that GFN1-xTB approximates MOF band gaps well, as compared to semilocal DFT. These data are used to train an interpretable Δ-learning model that predicts the difference between low- and high-fidelity band gaps, given by xTB and DFT data at the hybrid level, respectively. This model outperforms direct models trained using only the DFT values. The Δ-learning model also outperforms models with deep-learning architecture, fine-tuned on our custom dataset to predict the band gaps of MOFs. With limited high-quality DFT band gaps, taking advantage of Δ-learning using low-cost GFN1-xTB leads to better predictions than relying on DFT data alone.
A. Jose, A. Walsh· Journal of Chemical Theory a...· 0 citations
Melting point (MP) is an important thermophysical property for the chemical process industry, yet accurate prediction of MP for organic compounds in the absence of experimental data remains challenging due to the complex interplay between molecular packing, intermolecular interactions, and electronic structure. Traditional group contribution and quantitative structure-property relationship models, which rely primarily on static molecular descriptors, often fail to capture these critical condensed-phase effects. In this study, we present a hybrid machine learning framework that integrates cheminformatics descriptors with quantum chemical features and dynamic condensed-phase descriptors derived from molecular dynamics (MD) simulations. Using a curated subset of the DIPPR 801 database, multiple machine learning architectures, including light gradient boosting machine (LightGBM) and graph convolutional networks, were evaluated with feature sets of increasing physical fidelity. The best-performing model, based on LightGBM trained on Dragon descriptors augmented with MD and quantum chemical features, achieves a mean absolute error of 22.5 K, outperforming descriptor-only models and structure-based deep learning baselines. Shapley additive explanations interpretability analysis reveals that melting behavior is governed primarily by molecular topology, surface-area-weighted electronic descriptors, and condensed-phase interaction properties. In contrast, many isolated functional group and single molecule electronic descriptors contribute negligibly once these effects are accounted for. These results demonstrate that incorporating physics-informed, multi-scale descriptors enables more accurate and physically interpretable MP predictions.
Frank T. Mtetwa, N. Giles, W. Wilding et al.· Journal of Chemical Physics· 0 citations
Fluorinated molecules, including per- and polyfluoroalkyl substances (PFAS), present persistent challenges for thermochemical characterization due to limited experimental data, strong carbon–fluorine bonding, and the rapidly expanding size and diversity of fluorinated chemical space. While density functional theory (DFT) calculations can provide useful thermochemical data for individual fluorinated species, their routine application becomes increasingly impractical as molecular size, conformational complexity, and the number of distinct PFAS compounds continue to grow. Existing Benson-type group additivity schemes provide limited resolution for fluorinated environments, restricting their applicability to modern fluorinated and PFAS-relevant systems. Here, we develop a chemically resolved group additivity (GA) framework for fluorinated and PFAS-relevant species by fragmenting DFT-derived thermochemistry for 3070 molecules. This approach expands the available fluorinated Benson-type group library from 14 to 159 local environments and integrates the resulting groups within the Python Group Additivity (pGrAdd) framework. 10-fold cross-validated regression against DFT data yields root-mean-square deviations (RMSDs) of 8.14 kcal·mol–1 for enthalpy and 10.05 cal·mol–1·K–1 for entropy, which are reduced to 2.55 kcal·mol–1 and 6.36 cal·mol–1·K–1, respectively, following application of independently defined nongroup interaction correction terms in pGrAdd. Comparison with available experimental thermochemical data shows improved agreement and reduced bias compared to legacy Benson group libraries. This expanded fluorinated GA framework enables scalable and chemically interpretable thermochemical predictions for fluorinated and PFAS-relevant species, supporting kinetic modeling and mechanistic studies where direct electronic structure calculations are feasible but not scalable.
Samuel Eccles, Steven Pellizzeri· Journal of Physical Chemistr...· 0 citations
Proton-coupled electron transfer (PCET) mediated by hydroquinone and related molecules is key to natural and artificial energy conversion. The reactivity of these molecules depends on their bond dissociation free energy (BDFE), but studying the relationship between structure and thermochemistry across this chemical space has been limited by challenging experimental setup and high computational expense. Here, we present the first use of the AIMNet2 neural network potential to calculate average BDFE (BDFEavg) values for the 2H+/2e− dehydrogenation of about 200 000 hydroquinone-like compounds, including vicinal diamines, diols, and dithiols. Benchmarking against DFT calculations for 168 substituted ortho-phenylenediamines (opda) shows good agreement (R2 ∼ 0.84). Our analysis finds that the BDFEavg of diamines ranges from 50 to 80 kcal mol−1 and can be systematically tuned by modifying the backbone and N-substitution: electron-withdrawing groups raise BDFEavg by up to 15 kcal mol−1, while lower aromaticity in furan and thiophene backbones decreases BDFEavg by approximately 10 kcal mol−1 compared to the phenyl systems (∼65 kcal mol−1). Validation through cyclic voltammetry and reactivity studies with quinone oxidants for selected compounds supports the computational results. This extensive thermochemical database and a web-based prediction tool developed as a result of this work will offer valuable resources for designing PCET reagents for catalysis, energy storage, and biomedical uses.
Rajdeep Sarma, Yiwen Wang, David D Hebert et al.· Chemical Science· 0 citations
Covalent organic frameworks (COFs) are highly ordered, porous organic materials whose reticular construction from tailored nodes and linkers enables atomic-level control over structure and function. The design space of COFs is vast with virtually unlimited combinations of nodes, linkers, and functional groups. Interpretable machine learning (ML) offers a pathway to navigate this complexity by identifying the structural features that govern materials performance, yet interpretability often comes at the cost of predictive accuracy. In this work, we introduce a novel multiple-kernel learning framework that achieves both accuracy and mechanistic insight. A multiple-kernel ridge regression (MKRR) model was trained on band gaps predicted from GFN1-xTB level theory for a data set of 232 theoretical pyranoazacoronene (PAC) COFs produced from eight different conjugated linkers and 29 functional groups. Modifying these building units alone produced a range of band gaps between 0.4–2 eV. Manual analysis of the theoretical band gaps versus the linker indicates that breaking the conjugation pathway by altering the bond angle or by introducing a σ-bond increases the band gap while increasing the length of the linker decreases the band gap. All functional groups appear to reduce the band gap with three specific electron withdrawing groups reducing the band gap near 0.4 eV. For the ML, the building units were represented with three independent kernels that encoded the local environments of each node, linker, and functional group calculated from the Smooth Overlap of Atomic Positions (SOAP). After decomposing each kernel’s contribution to the model’s global predictions, we found that the MKRR model successfully captures the underlying structure–property relationships that influence the band gap. These results demonstrate that MKRR is an effective and interpretable framework for understanding and designing functional COFs.
Alathea E. Davies, O. Adesina, Isabella M. Valdez et al.· Journal of Chemical Theory a...· 1 citation