Page 70 - Read Online
P. 70
Tang et al. J. Mater. Inf. 2025, 5, 38 https://dx.doi.org/10.20517/jmi.2025.05 Page 11 of 19
data for training and test data points. The RMSE values are summarized in Table 3.
Figures 6-8 show the shapes of the component functions f(x) in the order of their importance for some of
i
i
the datasets selected to illustrate a case of a data subset with linear component functions, a case with non-
linear component functions, and the functions for the combined dataset. The shapes of the component
functions for the other datasets are shown in the Supplementary Figures 2-4). Their importance [evaluated
as a square root of the variance of f(x)] is shown in Figure 9.
i
i
The following can be concluded from these results. Optimal hyperparameters are quite different among the
subsets. This, coupled with data distributions indicating only partially overlapping volumes of sampled
space [Figures 2-4], explains difficulties of obtaining a better model by combining the subsets that were also
observed with SVR. It might instead be advisable to use separate ML models that are used depending on the
type of dataset (a category). The difference between the optimal length parameter (l) corresponds to
different roles played by nonlinearity. In particular, GPR-NN reveals that the dependence on the features is
practically linear for the Nb alloys dataset [Figure 6], with any noticeable nonlinearity appearing only in
f(x) whose contributions are minor [Figure 9]. For the datasets corresponding to Nb-Nb Si alloys,
i
i
3
5
nonlinearity is substantial [Figure 7 and Supplementary Figures 2-4]. It is very pronounced in the combined
dataset [Figure 8] as the algorithm attempts to learn a heterogeneous dataset leading to a small length
parameter [Table 3].
The relative feature importance is different between the datasets. Feature importance is known to be
method-dependent. Here it is clearly data-dependent and is different not just between Nb and Nb-Nb Si
3
5
alloys but also between Nb-Nb Si alloys alloyed at different substitution pair sites. Each vertical set of points
3
5
in Figure 9 shows the spread of feature importance values over 100 runs differing by random train-test split.
The existence of a substantial spread indicates a relatively sparse-data regime, but a relative persistence of
relative feature importance with different train-test splits indicates a degree of reliability of ML.
Nevertheless, we would caution against reading too much into feature importances obtained with black-box
algorithms as they are algorithm-dependent and may return nonsensical results (see Ref. for a spectacular
[52]
example of significant importance attached by ML to features which are numeric IDs of molecular blocks
devoid of any physical meaning).
Importantly, when adding coupling terms [i.e., Equation (6) with increasing N] there is little statistically
significant improvement in RMSE of test dataset for any of the subsets or for the combined set [Table 4]. In
some of the subsets there was only improvement in the training set error. This indicates that while
nonlinearity is important, inter-feature coupling terms are unimportant or non-recoverable probably due to
low sampling density in this case.
Analysis of feature dimensionality reduction
We performed uniform manifold approximation and projection (UMAP) dimensionality reduction of the
features and projection of substitution energies predicted via the CE-SVR models on the various datasets
[Figure 10]. The UMAP feature analyses show that the two feature vectors after dimensionality reduction
were not able to distinguish the target properties unambiguously in most datasets except for Nb and
Si -Nb Si with fewer data. These indicate the highly nonlinear relationships between the features and target
5
I
3
properties. Moreover, the distribution patterns with such non-linear feature-property relationships differ
significantly among various datasets. The various distribution patterns in the feature maps demonstrate that
the CE features indeed capture the structural variations of different substitution sites while they have the
same composition associated with the same substitutional elements. It also suggests that it is necessary to

