Page 62 - Read Online
P. 62
Tang et al. J. Mater. Inf. 2025, 5, 38 https://dx.doi.org/10.20517/jmi.2025.05 Page 3 of 19
Dataset augmentation and combination is another approach to palliate data scarcity. It can be done by
combining datasets from several related systems. For example, when using ML to predict the screening
factor of the SoftBV approximation [28,29] , data scarcity prevented effective ML for specific crystal structure
types - perovskite and spinel oxides - considered individually, whereby individual datasets only had about
100 data points. However, combining perovskite and spinel sets increased the accuracy of prediction for
both. Transfer learning methods can also be employed, enabling ML models that are initially trained on
datasets of spinel oxides to effectively predict the stability of perovskite oxides . Individual datasets may
[30]
help achieve a denser internal sampling of similar overlapping regions of feature space, but they can also
expand the volume of feature space. In the former case, and if the similarity of sampled systems translates
into hyperparameter similarity between datasets, their combination is likely to facilitate building a more
accurate ML model. If different datasets sample distinct parts of the feature space and/or require varying
hyperparameters, such data combinations may instead complicate building an accurate ML model. In this
work, the effect of data combination is discussed on the example of ML substitution energies of Nb and Nb-
Si alloys.
o
NbSi-based alloys are intermetallic compounds with high melting points (2,400 C) and low densities (6.6-
7.2 g/cm ), making them promising ultra-high temperature materials for next-generation aeroengine
3
[31]
turbines beyond current nickel-based superalloys . The Nb-Nb Si composites exhibit an excellent
3
5
combination of the toughness of Nb and the strength of Nb Si but suffer from problems with the strength/
3
5
oxidation resistance of Nb and the deformability of Nb Si . Many costly experimental works have shown
5
3
that adding alloying elements is an effective way to improve the comprehensive performance of Nb-Si
alloys [32,33] . DFT-based first-principles calculations have been used to study the stability and mechanical
properties of NbSi-based superalloys doped with various alloying elements. First-principles calculations are
also time-consuming, so only a very limited number of alloying elements and substitution sites have been
studied [34-36] . ML, as an emerging data-driven research paradigm in materials science, has proven to be
effective and efficient in describing complex structure-property relationships in materials [5,37,38] .
In this work, we therefore explore the effects of dataset combination and of feature nonlinearity and
coupling on the functional form of the dependence of the target property on the features in the optimal ML
model by considering the problem of prediction of substitution energies of Nb and Nb-Nb Si eutectic alloys
3
5
as a function of alloying elements and their substitution sites in Nb and Nb Si phases. In these systems, the
5
3
data can be naturally divided into subsets based on the type of substitutional site: the substitutional sites in
pure Nb and various inequivalent paired dual substitution sites in Nb and Nb-Nb Si phases. These sub-
5
3
datasets can be machine-learned separately or combined in a single model to examine the effect of data
combination. We use center-environment (CE) features [39,40] (defined below in more detail) that encode
chemical composition and structure information by using properties of constituent atoms projected onto a
composition- and structure-dependent basis set. While the basis set used for the projection is system-
dependent, the CE definition is generally applicable to any system, even including those with low local
symmetry. CE features have been successfully used to machine-learn various properties including structural
parameters, formation energies, band gaps, and molecular adsorption energies [30,39,40] . We perform ML with
common kernel methods, as well as with the Gaussian process regression-neural network (GPR-NN) hybrid
ML method that allows disambiguating the effects of feature nonlinearity and inter-feature coupling with
additive models. We find that different data subsets, corresponding to different substitution sites, expand
the volume of the feature space rather than just increase its internal sampling density, so that their
combination complicates rather than facilitates ML. Data for Nb alloys and Nb-Si alloys require different
optimal hyperparameters; while the best model for Nb alloys is practically linear, nonlinearity is important
for Nb-Si alloys, and it is more pronounced when ML the combined full dataset. We also find that while

