Author

Daisuke Yokogawa

1 paper indexed here

Fetches their full publication history.

Not the right person? Other researchers publish under this name.

Jul 2026

DFTB-Based Validation of Molecules Generated by Reversible Junction Tree Reinforcement Learning: A Case Study on Molecular Solar Thermal Fuels

Machine-learning methods are widely used for molecular generation across vast chemical spaces. However, most studies evaluate model performance using empirical descriptors such as logP and molecular weight, which do not indicate whether generated molecules satisfy target properties. Moreover, when generative models propose molecules outside existing databases—as is desirable in exploratory design—reference data are unavailable, making validation difficult. To move data-driven molecular design toward practical applications, it is essential to couple molecular generation with quantum-chemical calculations and to establish a validation workflow that screens many candidates while retaining physical reliability. Density-functional tight-binding (DFTB) balances accuracy and cost and is suitable for large-scale quantum-chemical screening. Reversible junction tree reinforcement learning (RJT-RL) generates molecules by assembling fragments on a reversible tree representation, combining interpretability with goal-directed optimization. In this work, we construct a validation framework that couples RJT-RL with DFTB calculations. Using azobenzene-based molecular solar thermal fuels (STFs) as a model system, we evaluate the property distributions of RL-generated molecules and compare them with reference molecules. On the generation side, we fix the azobenzene backbone and construct candidates by attaching substituents to the aromatic rings and, when applicable, further functionalizing them. A database containing approximately 5×10 4 azobenzene derivatives is used as an expert dataset to pretrain the RJT-RL model, allowing the policy network to learn structural patterns. In the subsequent reinforcement-learning stage, the reward is defined as the maximum Tanimoto similarity between each generated molecule and molecules in this database. Figure 1(a) shows that the maximum similarity initially fluctuates and then reaches a plateau. We therefore define the first 1–12k generated molecules as the oscillation stage and the 15–27k molecules as the platform stage, and perform DFTB calculations for all molecules generated in these two stages to compare generation behavior and properties before and after convergence of the reward. On the validation side, we apply a unified DFTB workflow to both RL-generated molecules and database molecules. SMILES strings are converted into three-dimensional trans and cis conformers using Open Babel. Geometry optimizations are then performed with DFTB+, followed by ground-state to first excited-state (S 0 →S 1 ) excitation-energy calculations on the optimized trans structures. This yields the excitation energy ΔE exc , relevant to matching the solar spectrum, and the energy difference ΔE iso between the trans and cis isomers, characterizing the energy-storage capacity. Candidates for which either geometry optimization or the excitation-energy calculation fails are excluded from the subsequent property statistics, and the overall DFTB success rate is used as a simple proxy for the structural reasonableness of the generated molecules. Successfully calculated molecules Success rate Oscillation stage (1-12k) 3100 25.8% Platform stage (15k-27k) 7992 66.6% This table summarizes the DFTB success rates in the two training stages. In the oscillation stage, DFTB calculations succeed for only 25.8% of generated molecules, whereas in the platform stage the success rate rises to 66.6%. This indicates that, under a similarity-based reward, the trained network produces a larger fraction of geometries that lie within the applicability domain of the DFTB model. Figure 1(b) compares the DFTB property distributions of reference database molecules and of molecules generated in the oscillation and platform stages. For ΔE iso , the overall distributions in the two training stages are similar to that of the database, suggesting that RL-generated molecules do not strongly deviate from the reference set in terms of energy-storage capacity. In contrast, the ΔE exc distribution of molecules from the platform stage is more concentrated than that from the oscillation stage and exhibits a pronounced peak in the energy region overlapping with the visible spectrum, indicating an enrichment of candidates in this desirable range compared with the reference database. Applying simple STF criteria, such as requiring ΔE exc to fall within a target window and ΔE iso to exceed a threshold, yields a subset of RL-generated molecules with potential relevance. At the same time, the weak correlation between the similarity-based reward and DFTB-computed ΔE exc and ΔE iso shows that the reward magnitude alone is insufficient to reliably predict property quality, highlighting the need for quantum-chemical validation of RL-generated molecules. In summary, we present a workflow that combines reversible junction tree reinforcement learning with rapid DFTB calculations, and demonstrate its use in assessing the chemical reasonableness and property distributions of similarity-driven RL-generated molecules in an azobenzene-based STF system. The framework is not limited to STF chemistry or specific target properties: by changing the reward definition and the set of properties computed with DFTB, it can be extended to other molecular design tasks, providing a general strategy for evaluating and calibrating the physical reliability of AI-generated molecules. Figure 1

Qingyu Zhu, Ryan T. Lingg, Ian McCleary et al. · 0 citations