Skip to content
Open access

GFN1-xTB-Assisted Machine Learning for Electronic Structure Screening of Metal–Organic Frameworks

Aug 2026 · Journal of Chemical Theory and Computation · 0 citations · 87 references

TL;DR

A semiempirical extended tight-binding approach (GFN1-xTB) is employed to compute the electronic properties of a dataset of MOFs, and it is shown that GFN1-xTB approximates MOF band gaps well, as compared to semilocal DFT.

Abstract

Metal–organic frameworks (MOFs) are versatile materials with tunable crystal structures, morphologies, and chemistries, offering diverse physical and chemical properties. Although typically electrically insulating, specific combinations of organic and inorganic components can impart electrical conductivity to MOFs. The virtually limitless chemical space of MOFs, however, presents a significant challenge in identifying optimal candidates for various applications. Although density functional theory (DFT) can probe their electronic structure, its high computational cost hinders the discovery of novel electroactive MOFs using machine learning due to limited data. To tackle these challenges, a semiempirical extended tight-binding approach (GFN1-xTB) is employed to compute the electronic properties of a dataset of MOFs, and it is shown that GFN1-xTB approximates MOF band gaps well, as compared to semilocal DFT. These data are used to train an interpretable Δ-learning model that predicts the difference between low- and high-fidelity band gaps, given by xTB and DFT data at the hybrid level, respectively. This model outperforms direct models trained using only the DFT values. The Δ-learning model also outperforms models with deep-learning architecture, fine-tuned on our custom dataset to predict the band gaps of MOFs. With limited high-quality DFT band gaps, taking advantage of Δ-learning using low-cost GFN1-xTB leads to better predictions than relying on DFT data alone.

Read PDF

Similar papers

Aug 2026

Δ -Machine Learning for the Prediction of Metal Complex Properties.

The discovery and design of novel transition metal complexes for specific applications heavily rely on computational high-throughput screenings to identify promising candidates for experimental validation. However, traditional computational approaches, such as density functional theory, are often too computationally demanding to be applied on a large scale. Machine learning methods offer a promising alternative due to their excellent computational efficiency, but their accuracy and high data requirements remain major challenges for their effective implementation. To address these issues, we herein present an adaptation of the Δ-ML strategy for quantum property prediction of transition metal complexes. We combine GFN2-xTB geometry optimizations and density functional theory single-point calculations in order to obtain low-fidelity approximations and generate featurized graph representations that serve as input to a graph neural network architecture. The high-fidelity targets originate from the tmQMg dataset and include the electronic and dispersion energies, HOMO-LUMO gap and dipole moment at the PBE0-D3BJ/def2-TZVP level as well as the polarizability at the PBE-D3BJ/def2-SVP level. Compared to a conventional benchmark approach, the proposed method consistently achieves higher accuracy in the prediction of high-fidelity targets, while demonstrating improved data efficiency and out-of-domain transferability. We furthermore show, how the use of cheaper low-fidelity methods leads to significant reductions in computational cost at minor losses in predictive performance. Overall, these results highlight the potential of Δ-ML for materials discovery in transition metal chemistry, which requires high predictive accuracy despite often times limited availability of training data.

Hannes Kneiding, David Balcells · 0 citations
Aug 2026

Multiple-Kernel Ridge Regression for Learning the Structure-Electronic Property Relationships of Pyranoazacoronene COFs

Covalent organic frameworks (COFs) are highly ordered, porous organic materials whose reticular construction from tailored nodes and linkers enables atomic-level control over structure and function. The design space of COFs is vast with virtually unlimited combinations of nodes, linkers, and functional groups. Interpretable machine learning (ML) offers a pathway to navigate this complexity by identifying the structural features that govern materials performance, yet interpretability often comes at the cost of predictive accuracy. In this work, we introduce a novel multiple-kernel learning framework that achieves both accuracy and mechanistic insight. A multiple-kernel ridge regression (MKRR) model was trained on band gaps predicted from GFN1-xTB level theory for a data set of 232 theoretical pyranoazacoronene (PAC) COFs produced from eight different conjugated linkers and 29 functional groups. Modifying these building units alone produced a range of band gaps between 0.4–2 eV. Manual analysis of the theoretical band gaps versus the linker indicates that breaking the conjugation pathway by altering the bond angle or by introducing a σ-bond increases the band gap while increasing the length of the linker decreases the band gap. All functional groups appear to reduce the band gap with three specific electron withdrawing groups reducing the band gap near 0.4 eV. For the ML, the building units were represented with three independent kernels that encoded the local environments of each node, linker, and functional group calculated from the Smooth Overlap of Atomic Positions (SOAP). After decomposing each kernel’s contribution to the model’s global predictions, we found that the MKRR model successfully captures the underlying structure–property relationships that influence the band gap. These results demonstrate that MKRR is an effective and interpretable framework for understanding and designing functional COFs.

Alathea E. Davies, O. Adesina, Isabella M. Valdez et al. · 1 citation
Preprint Jul 2026

Data-driven Design of Metal-Organic Frameworks with Tunable Negative Thermal Expansion

Materials with negative thermal expansion (NTE) are essential for applications requiring precise control of thermal expansion. Owing to their exceptional chemical tunability, flexible architectures, and low-energy lattice vibrations, metal-organic frameworks (MOFs) represent a rich platform for exploring NTE. However, uncovering the structural motifs that govern NTE across the enormous MOF design space remains experimentally challenging, and large-scale first-principles phonon calculations are computationally prohibitive. Here, we comprehensively evaluate the factors influencing NTE in MOFs by utilizing a high-throughput workflow based on MACE-MP-MOF0, a machine learning interatomic potential fine-tuned for MOFs with near-ab initio accuracy, to construct PhononMOFdb, a database of phonons, inelastic neutron scattering spectra, bulk moduli, and heat capacities for over 12,000 MOFs. High-throughput screening of this database reveals that highly porous cubic topology frameworks with heavier, lower-valent metal nodes favor strong NTE, while linker functionalization provides a practical handle for tuning NTE magnitude and sign without compromising mechanical stability. Experimental validation via high-resolution temperature-dependent synchrotron powder X-ray diffraction on the Ce-UiO-66 MOF and its brominated variants confirms the design recipe and yields volumetric NTE coefficients surpassing current records. This work establishes a data-driven strategy for engineering NTE in MOFs, showing how machine learning-accelerated discovery and targeted experimental validation together unlock predictive materials design.

Prathami Divakar Kamath, F. Tavani, A. Elena et al. · 0 citations
#explainable ai Review Aug 2026

Computational and ML methods in MOF based supercapacitors - from mechanistic understanding to future materials design

Metal organic frameworks (MOFs) have emerged as promising electrode materials for supercapacitor (SC) due to their high surface areas, tunable porosity, and redox active sites. However, the vast chemical space of MOFs leads to millions of possible structures, makes experimental trial and error discovery inefficient. This review provides a focused perspective on how density functional theory (DFT) and machine learning (ML) are enabling the accelerated discovery and rational design of MOF-based SC electrodes. Key insights from DFT are discussed in relation to three critical performance descriptors as electrical conductivity, electrochemical and structural stability, and redox activity. In parallel, recent advances in ML-driven screening are reviewed, covering the development and use of large-scale MOF databases, descriptor engineering strategies, and predictive model architectures. Case studies demonstrating the successful integration of DFT and ML for identifying high-performance MOFs are highlighted in the review. Finally, the current limitations are analysed, including the discrepancy between idealized computational models and real polycrystalline electrodes, intrinsic trade-offs between conductivity and stability, and the need for interpretable and physics-informed ML models. Overall, this review outlines a computational roadmap for the rapid discovery and optimization of next-generation MOF-based SC electrodes. This is the comprehensive review on ML-driven prediction of electrochemical performance in MOF based electrode for SC applications. It integrates DFT-calculated electronic descriptors with ML models to discover hidden structure-property relationships. It identifies critical data gaps, model transferability issues, lack of dynamic ion-transport modelling in current studies. It proposed a multi-fidelity active learning framework combining DFT, ML and experiments for accelerated MOF discovery. It outlines standardized database protocols and explainable AI strategies to guide future high-performance MOF design.

Achal Siddharth Fulmali, H. Panda · 0 citations
Jul 2026

tmGNN-XAI: An Explainable Graph Neural Network Tool for Predicting Electronic Properties of Transition Metal Complexes from SMILES.

Predicting the electronic properties of transition metal complexes (TMCs) from 2D molecular graphs remains challenging; organic-trained property models lack TMC transferability, universal interatomic potentials require 3D coordinates rather than SMILES, and tools providing holistic electronic property prediction with atom-level explainability and calibrated uncertainty remain limited. We present tmGNN-XAI, a multitask relational graph convolutional network that predicts seven quantum-chemical properties of TMCs directly from SMILES strings and produces perturbation-based atom-level attributions for each prediction. The model encodes dative coordination bonds as a dedicated edge type distinct from covalent bonds and is trained on 100,703 complexes from the tmQM data set spanning 30 transition metals. Test-set performance is competitive with a Chemprop D-MPNN baseline, achieving R2 = 0.979 for metal partial charge and R2 = 0.964 and 0.949 for HOMO and LUMO energies. Across all 100,703 complexes, donor atoms (N, O, S, P) appear among the top-five most important atoms in more than 99.8% of complexes for every property, a large-scale data-driven result consistent with ligand field theory. A trust framework combining ensemble agreement with attribution direction separates predictions into four reliability scenarios; confident predictions achieve 1.6 to 2.5 times lower mean absolute error than uncertain ones for five of seven properties. The framework generalizes to cross-level DFT validation, phototherapy candidate screening (area under the ROC curve (AUC) = 0.735), and indirect redox prediction via Koopmans' theorem. An interactive web application makes property predictions, atom-level attributions, and trust labels accessible without programming or DFT expertise. tmGNN-XAI is designed as an explainable, first-tier screening tool for TMC electronic property estimation.

Abdulmujeeb T. Onawole · 2 citations
Preprint Jul 2026

Machine learning based prediction of optical properties in two-dimensional Mo-W-S-Se-Te transition-metal dichalcogenide alloys through physics-informed sampling

Two-dimensional transition-metal dichalcogenide (TMD) alloys provide a compositionally tunable platform for controlling the optical and electronic properties. However, systematic prediction of their dielectric response across multicomponent alloy spaces remains challenging owing to the combinatorial cost of first-principles calculations. In this work, we combined ab initio optical-property calculations with a tabular foundation-model regression to predict the real and imaginary components of the frequency-dependent dielectric function for Mo-W-S-Se-Te TMD alloys. A dataset of 99 alloy structures spanning binary, ternary, quaternary, and quinary compositions was generated using density functional theory (DFT). The resulting polarization-dependent dielectric spectra were used to train a tabular prior-fitted network (TabPFN) and evaluated against the conventional Extra Trees and XGBoost models. To accommodate the in-context capacity limit of TabPFN, we introduced a non-uniform, physics-informed energy subsampling strategy that concentrates sampling in the optically active region above the band gap, where interband absorption is strongest. Trained solely on quaternary alloys, our TabPFN reconstructed the dielectric spectra of held-out quaternary compositions with an R2>0.98 and a mean absolute error below 0.10 for all four dielectric components, outperforming both baselines while requiring no gradient-based training or hyperparameter tuning. Our model further predicted derived optical quantities, including refractive index, extinction coefficient, and absorption coefficient. Additionally, our model generalized in a zero-shot manner to binary, ternary, and quinary alloys absent from the training set, with quinary predictions achieving an R2>0.97.

Vivek Chowdhury, Tarvir Anjum Aditto, M. Samrat et al. · 0 citations