Page 70 - Read Online
P. 70

Tang et al. J. Mater. Inf. 2025, 5, 38  https://dx.doi.org/10.20517/jmi.2025.05  Page 11 of 19

               data for training and test data points. The RMSE values are summarized in Table 3.


               Figures 6-8 show the shapes of the component functions f(x) in the order of their importance for some of
                                                                   i
                                                                i
               the datasets selected to illustrate a case of a data subset with linear component functions, a case with non-
               linear component functions, and the functions for the combined dataset. The shapes of the component
               functions for the other datasets are shown in the Supplementary Figures 2-4). Their importance [evaluated
               as a square root of the variance of f(x)] is shown in Figure 9.
                                              i
                                            i
               The following can be concluded from these results. Optimal hyperparameters are quite different among the
               subsets. This, coupled with data distributions indicating only partially overlapping volumes of sampled
               space [Figures 2-4], explains difficulties of obtaining a better model by combining the subsets that were also
               observed with SVR. It might instead be advisable to use separate ML models that are used depending on the
               type of dataset (a category). The difference between the optimal length parameter (l) corresponds to
               different roles played by nonlinearity. In particular, GPR-NN reveals that the dependence on the features is
               practically linear for the Nb alloys dataset [Figure 6], with any noticeable nonlinearity appearing only in
               f(x) whose contributions are minor [Figure 9]. For the datasets corresponding to Nb-Nb Si  alloys,
                  i
                i
                                                                                                   3
                                                                                                 5
               nonlinearity is substantial [Figure 7 and Supplementary Figures 2-4]. It is very pronounced in the combined
               dataset [Figure 8] as the algorithm attempts to learn a heterogeneous dataset leading to a small length
               parameter [Table 3].
               The relative feature importance is different between the datasets. Feature importance is known to be
               method-dependent. Here it is clearly data-dependent and is different not just between Nb and Nb-Nb Si
                                                                                                         3
                                                                                                       5
               alloys but also between Nb-Nb Si  alloys alloyed at different substitution pair sites. Each vertical set of points
                                           3
                                         5
               in Figure 9 shows the spread of feature importance values over 100 runs differing by random train-test split.
               The existence of a substantial spread indicates a relatively sparse-data regime, but a relative persistence of
               relative feature importance with different train-test splits indicates a degree of reliability of ML.
               Nevertheless, we would caution against reading too much into feature importances obtained with black-box
               algorithms as they are algorithm-dependent and may return nonsensical results (see Ref.  for a spectacular
                                                                                          [52]
               example of significant importance attached by ML to features which are numeric IDs of molecular blocks
               devoid of any physical meaning).


               Importantly, when adding coupling terms [i.e., Equation (6) with increasing N] there is little statistically
               significant improvement in RMSE of test dataset for any of the subsets or for the combined set [Table 4]. In
               some of the subsets there was only improvement in the training set error. This indicates that while
               nonlinearity is important, inter-feature coupling terms are unimportant or non-recoverable probably due to
               low sampling density in this case.

               Analysis of feature dimensionality reduction
               We performed uniform manifold approximation and projection (UMAP) dimensionality reduction of the
               features and projection of substitution energies predicted via the CE-SVR models on the various datasets
               [Figure 10]. The UMAP feature analyses show that the two feature vectors after dimensionality reduction
               were not able to distinguish the target properties unambiguously in most datasets except for Nb and
               Si -Nb Si  with fewer data. These indicate the highly nonlinear relationships between the features and target
                    5
                 I
                       3
               properties. Moreover, the distribution patterns with such non-linear feature-property relationships differ
               significantly among various datasets. The various distribution patterns in the feature maps demonstrate that
               the CE features indeed capture the structural variations of different substitution sites while they have the
               same composition associated with the same substitutional elements. It also suggests that it is necessary to
   65   66   67   68   69   70   71   72   73   74   75