Page 216 - Read Online
P. 216
Wang et al. J. Mater. Inf. 2026, 6, 16 Page 19 of 25
OUTLOOK: CHALLENGES AND OPPORTUNITIES
This section outlines both the challenges that constrain current progress and the opportunities that may
define the future of MF learning in materials design.
Despite notable progress, several fundamental challenges remain. First, data scarcity and imbalance are
persistent barriers. HF data, whether from advanced simulations or experiments, remain expensive and
unevenly distributed across material classes. The absence of standardized benchmarks further complicates
fair comparison among methodologies, slowing methodological innovation and adoption.
High dimensionality further exacerbates effective data scarcity in MF learning. As the number of
compositional, structural, and processing descriptors grows, the design space expands rapidly, increasing the
amount of data required for reliable surrogate modeling. This challenge is particularly pronounced for MF
methods, where accurate learning of LF-HF correlations often relies on co-located observations in
high-dimensional spaces. As a result, simply adding more LF data may yield only limited benefits. This
motivates MF strategies that aim to reduce the effective dimensionality of the problem, for example, through
physics-informed feature engineering, active-subspace or manifold identification, and learned latent
representations.
Beyond data volume and dimensionality, incomplete or partially observed feature sets also induce an implicit
degradation of data fidelity. When key compositional, microstructural, processing, or environmental
descriptors are missing, physically distinct material states may collapse onto similar representations in
feature space, inflating epistemic uncertainty in poorly characterized regions. In heterogeneous materials
databases, records with more complete descriptor sets can therefore be regarded as higher-fidelity
observations, whereas samples with missing variables represent degraded-fidelity views that should be
treated explicitly rather than discarded. For example, Ching and Phoon employed a hierarchical Bayesian
framework with Gibbs sampling to infer missing feature values prior to MF fusion, demonstrating improved
robustness and predictive performance when only partial variables are available. This perspective highlights a
fundamental connection between data completeness and fidelity and motivates future MF frameworks that
jointly address missing-data inference and fidelity-aware learning .
[105]
Second, transferability across material classes and design domains is limited. Models trained on MF datasets
for alloys or oxides, for example, may not generalize to polymers or energy-storage materials without
extensive retraining.
This lack of universality reflects both intrinsic differences in physics across classes and the difficulty of
encoding fidelity relationships in a transferable manner.
Third, interpretability and trust in black-box methods remain major obstacles. Deep-learning frameworks
can achieve impressive accuracy in MF prediction tasks, but their opaque decision making makes it difficult
to diagnose when LF biases are propagating unchecked. Without greater transparency and interpretability,
the adoption of these methods in high-stakes applications, such as structural materials or energy
technologies, will remain limited.
A particularly salient theme is the integration of MF methods with automated laboratories and closed-loop
workflows. On the one hand, current approaches face significant barriers: to guide experiments in real time,

