Page 216 - Read Online
P. 216

Wang et al. J. Mater. Inf. 2026, 6, 16                                           Page 19 of 25




               OUTLOOK: CHALLENGES AND OPPORTUNITIES
               This section outlines both the challenges that constrain current progress and the opportunities that may
               define the future of MF learning in materials design.


               Despite notable progress, several fundamental challenges remain. First, data scarcity and imbalance are
               persistent barriers. HF data, whether from advanced simulations or experiments, remain expensive and
               unevenly distributed across material classes. The absence of standardized benchmarks further complicates
               fair comparison among methodologies, slowing methodological innovation and adoption.


               High dimensionality further exacerbates effective data scarcity in MF learning. As the number of
               compositional, structural, and processing descriptors grows, the design space expands rapidly, increasing the
               amount of data required for reliable surrogate modeling. This challenge is particularly pronounced for MF
               methods, where accurate learning of LF-HF correlations often relies on co-located observations in
               high-dimensional spaces. As a result, simply adding more LF data may yield only limited benefits. This
               motivates MF strategies that aim to reduce the effective dimensionality of the problem, for example, through
               physics-informed feature engineering, active-subspace or manifold identification, and learned latent
               representations.


               Beyond data volume and dimensionality, incomplete or partially observed feature sets also induce an implicit
               degradation of data fidelity. When key compositional, microstructural, processing, or environmental
               descriptors are missing, physically distinct material states may collapse onto similar representations in
               feature space, inflating epistemic uncertainty in poorly characterized regions. In heterogeneous materials
               databases, records with more complete descriptor sets can therefore be regarded as higher-fidelity
               observations, whereas samples with missing variables represent degraded-fidelity views that should be
               treated explicitly rather than discarded. For example, Ching and Phoon employed a hierarchical Bayesian
               framework with Gibbs sampling to infer missing feature values prior to MF fusion, demonstrating improved
               robustness and predictive performance when only partial variables are available. This perspective highlights a
               fundamental connection between data completeness and fidelity and motivates future MF frameworks that
               jointly address missing-data inference and fidelity-aware learning .
                                                                      [105]

               Second, transferability across material classes and design domains is limited. Models trained on MF datasets
               for alloys or oxides, for example, may not generalize to polymers or energy-storage materials without
               extensive retraining.


               This lack of universality reflects both intrinsic differences in physics across classes and the difficulty of
               encoding fidelity relationships in a transferable manner.

               Third, interpretability and trust in black-box methods remain major obstacles. Deep-learning frameworks
               can achieve impressive accuracy in MF prediction tasks, but their opaque decision making makes it difficult
               to diagnose when LF biases are propagating unchecked. Without greater transparency and interpretability,
               the adoption of these methods in high-stakes applications, such as structural materials or energy
               technologies, will remain limited.

               A particularly salient theme is the integration of MF methods with automated laboratories and closed-loop
               workflows. On the one hand, current approaches face significant barriers: to guide experiments in real time,
   211   212   213   214   215   216   217   218   219   220   221