Page 73 - Read Online
P. 73
Page 14 of 19 Tang et al. J. Mater. Inf. 2025, 5, 38 https://dx.doi.org/10.20517/jmi.2025.05
Figure 9. Feature importance in the additive model for (A) Nb alloys, (B) Nb -Nb Si alloys, (C) Nb -Nb Si alloys, (D) Si -Nb Si alloys,
I 5 3 II 5 3 I 5 3
(E) Si -Nb Si alloys, (F) combined data set. Blue points are for the training, and red points are for the test set. The distribution of
II 5 3
vertical points at each feature is over 100 runs differing by random training-test splits.
This allowed us to explore the effect on the quality of ML of using individual data sets for each type of
substitution or aggregating data corresponding to different types of substitutional sites. Under sparse data,
data combination is typically believed to be promising. We demonstrated here on the example of NbSi
alloys that data combination does not necessarily improve the quality of an ML model if the data do not
sample similar areas in the feature space. The prediction error for the combined dataset is higher than
prediction errors for subsets corresponding to individual different site types. This reflects data distribution,
namely that the data from subsets only partially overlap in the feature space and, rather than increasing the
density of internal sampling which would facilitate ML, increase the volume of feature space. In this case, it
may be recommended to use individual models for the subsets as the perceived advantages of a bigger
combined dataset may not be realized. The transferability issues should be paid attention to for doped
systems with different substitution configurations.

