Skip to content

Author

Liuqing Yang

2 papers indexed here

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

#small language model Dataset Open access Sep 2026

Evidence and Metadata Dataset for "Large Language Models for Small-Molecule Discovery: Model Taxonomy, Chemical Grounding, and Experimental Evidence

This dataset supports the review “Large Language Models for Small-Molecule Discovery: Model Taxonomy, Chemical Grounding, and Experimental Evidence”. It provides a structured, machine-readable evidence base for the literature synthesis, including: (i) a literature inventory and search/screening seed; (ii) model metadata for representative molecular language models, chemistry-specialized large language models, multimodal chemical foundation models, and integrated molecular-design systems; (iii) a study-level experimental evidence matrix distinguishing automated execution from iterative and adaptive experimental feedback; (iv) datasets and benchmark resources spanning molecular properties, reactions, 3D structures, molecule–text data, instruction resources, spectroscopy, and chemistry-LLM evaluation; (v) taxonomy and extraction codebooks; and (vi) source/provenance registries and validation scripts. The evidence cutoff is 4 September 2026. Unreported values are encoded as NR (“not reported with sufficient specificity”); no synthetic or idealized value is treated as reported evidence. The Stage-2 dataset contains 45 core studies with primary or official source locators, 21 structured model/system records, 8 experimental or closed-loop comparator records, 18 dataset/benchmark resources, and 68 provenance/source-registry entries. The dataset is designed to support reproducibility, evidence auditing, model taxonomy, and study-level assessment of experimental maturity in language-centred small-molecule discovery. A final frozen version will replace this draft before publication of the associated review.

Liuqing Yang, Zhongjian Wang, Hong Wan et al. · 0 citations
#small language model Dataset Open access Sep 2026

Evidence and Metadata Dataset for "Large Language Models for Small-Molecule Discovery: Model Taxonomy, Chemical Grounding, and Experimental Evidence

This dataset supports the review “Large Language Models for Small-Molecule Discovery: Model Taxonomy, Chemical Grounding, and Experimental Evidence”. It provides a structured, machine-readable evidence base for the literature synthesis, including: (i) a literature inventory and search/screening seed; (ii) model metadata for representative molecular language models, chemistry-specialized large language models, multimodal chemical foundation models, and integrated molecular-design systems; (iii) a study-level experimental evidence matrix distinguishing automated execution from iterative and adaptive experimental feedback; (iv) datasets and benchmark resources spanning molecular properties, reactions, 3D structures, molecule–text data, instruction resources, spectroscopy, and chemistry-LLM evaluation; (v) taxonomy and extraction codebooks; and (vi) source/provenance registries and validation scripts. The evidence cutoff is 4 September 2026. Unreported values are encoded as NR (“not reported with sufficient specificity”); no synthetic or idealized value is treated as reported evidence. The Stage-2 dataset contains 45 core studies with primary or official source locators, 21 structured model/system records, 8 experimental or closed-loop comparator records, 18 dataset/benchmark resources, and 68 provenance/source-registry entries. The dataset is designed to support reproducibility, evidence auditing, model taxonomy, and study-level assessment of experimental maturity in language-centred small-molecule discovery. A final frozen version will replace this draft before publication of the associated review.

Liuqing Yang, Zhongjian Wang, Hong Wan et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.