A materials property hierarchy is introduced, from intrinsic, composition-determined properties to extrinsic, processing-dependent performance, to clarify deployment constraints and distinguish structural, physical and deployment novelty.
Abstract
Artificial intelligence (AI) is accelerating materials prediction and design by enabling efficient exploration of chemical and structural spaces, with particular promise for novel materials discovery. However, novelty in materials discovery encompasses chemical plausibility, structural distinctiveness, property relevance and experimental realisability, making AI-driven novelty claims difficult to substantiate. We introduce a materials property hierarchy, from intrinsic, composition-determined properties to extrinsic, processing-dependent performance, to clarify deployment constraints and distinguish structural, physical and deployment novelty. This framework motivates an evidence-based view of multimodal materials data spanning chemical composition, microstructure, processing, and testing and characterisation, showing that current evidence remains concentrated in composition and idealised structure while heterogeneous, under-represented and weakly integrated modalities limit support for physical and deployment novelty. It also highlights the limitations of benchmarks based mainly on computational labels and proxy novelty criteria. Community-wide standards for data collection, modality alignment and evidence synthesis are needed to support multimodal data construction, process-aware multimodal modelling, feasibility-first generative modelling and deployment-aware benchmarking, so that generative and multimodal AI can design experimentally realisable materials with defensible scientific and practical novelty.
Data-driven material selection is progressively changing how materials are evaluated in engineering, manufacturing, and product design. With the growing diversity of heterogeneous data sources—ranging from microscopic images and physicochemical properties to simulation outputs and textual data—multimodal machine learning (MML) has become a key technology for fusing different types of information and supporting reliable multi-objective decision-making. However, despite increasing research interest, there is still no unified and conceptually structured review on the application of MML methods in data-driven material selection. This paper provides a thorough and structured overview of the field, organized using a novel multidimensional classification framework based on data types, integration levels, learning paradigms, and decision-making tasks. Guided by this framework, we critically evaluate the strengths, limitations, and general applicability of existing methods, and identify current trends and research gaps. Beyond qualitative synthesis, we quantify the cited corpus (N = 94) by publication year, method family, and application area, and collate an empirical fusion-evidence table that reports each study’s quantified gain together with its boundary conditions. We further analyze the main challenges and present a forward-looking research agenda covering self-supervised learning, knowledge-enhanced models, and interpretable human–AI collaboration. This review aims to offer a conceptual framework and a solid reference for building intelligent multimodal material selection systems.
Yuwei Zhang, Chou Yong Tan, Beichen Wang et al.· IEEE Access· 0 citations
Generative artificial intelligence is rapidly transforming materials design by enabling de novo exploration of immense chemical spaces. Yet a large proportion of AI-generated compositions remain implausible, violating established chemical principles, which limits the reliability and interpretability of generative materials design. Here, we introduce a chemical validity operator that recasts heuristic chemical rules as a configurable algorithmic prior for evaluating and guiding generative materials discovery. Built on the open-source SMACT package, a data-informed oxidation-state model exposes tunable thresholds, allowing users to interpolate continuously between permissive and conservative chemical constraints, while supporting both exploratory and conservative materials-design workflows. Benchmarking six state-of-the-art generative models for inorganic crystals shows that most reproduce stoichiometry but under-represent realistic oxidation-state combinations, and that filtering removes compositions reliant on rarely observed oxidation states while preserving low-energy compounds near the convex hull. Beyond screening, the same operator can also serve as a reinforcement-learning reward, steering a latent diffusion model towards chemically grounded compositions. By encoding chemical heuristics and observations, this work establishes a foundation for oxidation-state-aware generative models.
Kinga O. Mastej, Panyalak Detrattanawichai, Hyunsoo Park et al.· 0 citations
From first-principles calculations to machine learning-driven materials discovery, computational methods enhance our understanding of material behavior under different conditions. Furthermore, these modern computational tools allow scientists to explore vast design spaces more efficiently. Namely, by simulating material properties before synthesis, researchers can rapidly screen potential candidates, optimize structures, and uncover novel materials that might not have been feasible through traditional experimentation alone. As the power of computation continues to grow, the role of computational tools in solving complex materials science challenges will only expand, accelerating innovation and transforming the way materials are understood, discovered, and developed.
I will present examples from our research that illustrate how we integrate high-throughput computing with machine learning (ML) and artificial intelligence (AI) to tackle complex challenges in materials science [1, 2]. I will first discuss recent progress in the development of automated computational workflows that support large-scale screening of materials for targeted properties, such as high electro-conversion or stability against electrochemical dissolution. These frameworks also allow us to develop large databases of relevant materials properties. When combined with modern ML tools, these databases can be used to train surrogate ML-models capable of screening millions of candidate chemistries to identify the ones with optimal reactivity or stability. Furthermore, I will illustrate the use of interpretable ML in materials research that aims to enhance explainability of predictive ML models, enabling understanding of the underlying factors influencing design decisions. Lastly, I will present our current work on inverse material design, where AI methods—particularly generative pretrained transformers—are used to predict new material candidates based on desired properties, pushing the boundaries of materials innovation.
[1] M. Davis, W. Kort-Kamp, E. F. Holby, P. Zelenay, and I. Matanovic, Computational Screening of Transition Metal-Nitrogen-Carbon Materials as Electrocatalysts for CO
2
Reduction.
Electrochimica Acta 510,
145357(2025).
[2] M. Davis, W. Kort-Kamp, I. Matanovic, P. Zelenay and E. F. Holby, Design of Amine-Functionalized Materials for Direct Air Capture Using Integrated High-Throughput Calculations and Machine Learning, accepted in
Communications Chemistry
(2025).
I. Gonzales, R. Ullberg, Andrew H Salij et al.· ECS Meeting Abstracts· 0 citations
This review examines emerging AI methodologies for accelerated materials discovery, with particular emphasis on how computational design, data infrastructure, synthesis planning, and autonomous experimentation can be connected into experimentally grounded workflows.
Jaehwan Choi, Seongmin Kim, Junkil Park et al.· Chemical Reviews· 0 citations
A reinforcement learning framework that guides latent denoising diffusion models in finding diverse and novel, yet thermodynamically viable, crystalline compounds and demonstrates enhanced property-guided design that preserves chemical validity while targeting desired functional properties.
Hyunsoo Park, Aron Walsh· Nature Machine Intelligence· 27 citations· ⚡3
Composite materials design requires understanding complex microstructural characteristics, necessitating the integration of heterogeneous data sources with artificial intelligence. Current multimodal learning frameworks are mostly developed for crystalline or polymer systems with discrete structure-property mappings and well-defined structural graphs, but fail to model the continuous and nonlinear composite design spaces under data scarcity. Here we present ordinality as a core principle when building multimodal representations for composite materials. We propose ORDinal-aware imagE-tabulaR (ORDER) alignment to integrate microstructures with tabular material descriptors and apply physics-based surrogate signals to eliminate the need for full property annotation. ORDER ensures similar target properties occupy nearby regions in the latent space, preserving the continuous nature of composite properties and enabling meaningful interpolation between sparsely observed designs. Evaluated on nanofiber and carbon fiber composite datasets, ORDER consistently outperforms alignment-oriented and property-aware baselines across property prediction, cross-modal retrieval, and microstructure generation tasks. Composite materials design is difficult due to complex, continuous design space. This work builds ORDER, a multimodal framework linking microstructures and descriptors with preserved property trends, aiding prediction, retrieval, and microstructure generation.
Xinyao Li, Hangwei Qian, Jingjing Li et al.· Nature Communications· 0 citations