Page 61 - Read Online
P. 61

Page 2 of 19                        Tang et al. J. Mater. Inf. 2025, 5, 38  https://dx.doi.org/10.20517/jmi.2025.05

               hyperparameters. The Gaussian process regression-neural network hybrid ML method was used to separate the
               effects of nonlinearity and inter-feature coupling and show that while for Nb alloys nonlinearity is unimportant, it is
               critical to Nb-Nb Si  alloys. We find that inter-feature coupling terms are unimportant or non-recoverable,
                                3
                             5
               demonstrating the utility of more robust and interpretable additive models.
               Keywords: Materials informatics, machine learning, kernel regression, feature engineering, alloys



               INTRODUCTION
               Prediction of materials properties from descriptors of chemical composition and structure with machine
               learning (ML) - a major part of materials informatics - is attracting growing attention, as it holds the
               promise of more rapid discovery of novel functional materials with desired properties by reducing the
               amount of the required experimentation and/or direct calculations of the target properties of material
               candidates and the associated human, CPU, and time costs. This is an enticing proposition, and for this
               reason, materials informatics is rapidly becoming one of research mainstreams, with ML of various
               properties such as formation energetics, band structure-derived properties, interaction and reaction
               energies and barriers reported in many works . Building accurate ML models requires good training data,
                                                      [1-5]
               potentially large amounts of data with high quality. Training data [either experimental or computational,
               e.g., with density functional theory (DFT) ] are often expensive to obtain, so that in many materials
                                                     [6]
               informatics problems, one has to deal with rather small datasets [7-10] . The number of features or descriptors
               D used is often large, resulting in hard high-dimensional ML problems. For example, atomic properties
               such as nuclear charge, ionization potential (IP), electron affinity (EA), and atomic orbital data are always
               available and are often used, resulting in dozens of features even for materials with few types of constituent
               atoms [11-15] . To this are typically added a number of features describing composition (stoichiometry),
               structure, atoms or electron densities etc. [16-18] , resulting in potentially very high dimensional feature spaces.


               The density of sampling with a limited-size dataset in a high-dimensional feature space is thus bound to be
               low, and the data scarcity issues arise pertaining to the proverbial “curse of dimensionality” . When using
                                                                                             [19]
               common nonlinear ML methods, this manifests itself in overfitting on the one hand and in methodological
               issuefs on the other, e.g., loss of the advantage of using Matern kernels. One way to deal with it is to use
               more robust models including linear models or non-linear models - linear regressions using preset
               customized nonlinear basis functions, either those reflecting the nature of the underlying phenomena or
               those found to provide a good fit [20,21]  as opposed to, e.g., generic nonlinear kernels. Another way is to build
               the target function from component functions of lower dimensionality. The use of the latter has been
               formalized with high-dimensional model representation (HDMR) [22-24]  ideology. HMDR can be effectively
               done with ML [25-27]  but suffers from a combinatorial growth of the number of terms with both D and the
               included order of coupling among the features. However, simple additive models (1st order HDMR) do not
               suffer from this issue (as the number of terms then is simply D) and are attractive if inter-feature coupling is
               unimportant or unrecoverable due to low data density . These approaches allow interpretability as they
                                                              [25]
               help reveal the functional form of the dependence of the target function on features [20,21,26] ; this information
               can then be used to guide future model construction, either analytic or algorithmic. Pieces of information
               that are expected to be useful for this include the knowledge of whether the underlying functional
               dependence is linear or nonlinear and whether inter-feature coupling is important for a given practical
               problem. These issues are studied here in a case study of ML of substitution energies of alloying elements in
               Nb and Nb-Si alloys.
   56   57   58   59   60   61   62   63   64   65   66