Page 199 - Read Online
P. 199
Page 2 of 25 Wang et al. J. Mater. Inf. 2026, 6, 16
INTRODUCTION
Data-driven methodologies are increasingly being utilized to accelerate the design and discovery of new
materials [1-11] . At the core of these approaches lies data, which provide the foundation for training predictive
models and guiding design strategies [12-17] . Yet the accelerating pace of materials discovery is constrained by a
persistent data bottleneck, where high-fidelity (HF) data are scarce while low-fidelity (LF) data are
abundant [18-23] . To formalize this disparity in data quality, we introduce the notion of data fidelity. Here,
“fidelity” refers to the quality of a data source or model, encompassing its systematic bias relative to the
desired high-truth target, its noise level or predictive variance, and the completeness/integrity of the
recorded variables. HF sources, such as density functional theory (DFT) calculations at the hybrid functional
level [9,24-27] or experimental measurements of structural, electronic, and mechanical properties [28-30] , are
regarded as the gold standard. Although these methods provide accurate and reliable information, they are
costly in terms of computational resources, experimental effort, or wall-clock time, which limits their
scalability across the vast chemical and structural design space . By comparison, LF data, including
[31]
semi-empirical simulations, classical force fields, and simplified experiments, are much more accessible but
inevitably less accurate and possibly biased [28-30,32,33] . Bridging the divide between scarce, high-quality data and
plentiful, low-quality data remains a central challenge for materials informatics. This challenge has motivated
the development of strategies capable of jointly leveraging information across fidelity levels.
Among these strategies, multi-fidelity (MF) learning has emerged as a powerful approach to overcome such a
bottleneck. In essence, MF learning refers to a class of machine learning approaches that integrate
information from datasets of varying accuracy and cost, such as combining LF simulations with HF
experimental measurements [34-36] . The central idea is to exploit the complementary strengths of different data
sources: LF data provide broad coverage of the design space and capture general trends in materials
performance, while HF data anchor the models with reliable accuracy [28,37,38] . By learning correlations and
systematic discrepancies between fidelity levels, MF approaches can effectively transfer knowledge across
datasets, yielding predictions that are both accurate and data-efficient [34,39-45] . This paradigm is particularly
promising for materials design, where exhaustive HF characterization is prohibitively expensive, yet the
demand for predictive accuracy is high [28-30,32,33] . Key methodological ideas include statistical and parametric
models, machine learning models with fidelity features, correction-based models such as co-kriging, deep
learning frameworks, and active learning frameworks, each offering distinct trade-offs in terms of accuracy,
scalability, and interpretability [46-50] . Together, these strategies enable the construction of surrogate models
with higher accuracy and lower costs, thereby accelerating materials discovery [29,51-53] .
MF learning has broad practical relevance across materials research, spanning high-throughput screening
pipelines [42,54] , interatomic potential predictions [38-40] , and data-driven design of new materials [37,55] . In steel
design, it couples fast thermo-micromechanical surrogates with strategically chosen HF finite element
evaluations to optimize strain-hardening performance with both efficiency and accuracy . In halide
[53]
perovskites, MF machine learning integrates semi-local and hybrid-functional DFT data with experimental
measurements to predict decomposition energies, guiding the genetic algorithm-driven discovery of stable
photovoltaic alloys . In polymers, a MF co-training neural network fuses all-atom molecular dynamics
[56]
simulation data (LF) with scarce experimental measurements (HF) to accurately predict phthalonitrile
melting points, enabling low-cost screening of low-melting candidates for improved processability . In
[57]
geomaterials, an active-learning MF residual Gaussian process (GP) framework was trained on abundant LF
data collected across multiple sites, with predictive uncertainty used as an acquisition function to guide HF
sampling. This approach demonstrated robust extrapolation in sparsely sampled regions and significantly

