Page 63 - Read Online
P. 63
Page 4 of 19 Tang et al. J. Mater. Inf. 2025, 5, 38 https://dx.doi.org/10.20517/jmi.2025.05
feature nonlinearity is important, inter-feature coupling terms are unimportant or non-recoverable, in
individual data subsets and in the combined full dataset, demonstrating the feasible utility of more robust
and interpretable 1st order additive models.
MATERIALS AND METHODS
Data sets
The training dataset is constructed based on first-principles calculations of alloyed Nb and α-Nb Si . Nb has
3
5
a body-centered cubic (BCC) crystal structure, while α-Nb Si adopts a body-centered tetragonal (BCT)
5
3
structure. In the pure Nb supercell, all Nb atoms are equivalent due to their symmetrical nature. In contrast,
the conventional cell of α-Nb Si contains two inequivalent Nb sites (dubbed Nb and Nb ) and two
II
I
5
3
inequivalent Si sites (dubbed Si and Si ) that can be substituted with alloying elements. Figure 1 shows Nb
I
II
and α-Nb Si systems with substitution sites for alloying elements. See Ref. for a more detailed description
[41]
5
3
of the crystal structure and sites. This 32-atom conventional cell consists of 20 Nb atoms and 12 Si atoms,
with four Nb , 16 Nb , four Si , and eight Si atoms, respectively. By considering site substitutions at the
II
I
II
I
non-equivalent site pairs with 14 different alloying elements, including B, Al, Si, Ti, V, Cr, Fe, Co, Ni, Y, Zr,
Nb, Mo, and Hf, we compiled a total of 3,738 double-site substitution energies (E ), which includes 210
DS
data points for Nb and 3,528 for α-Nb Si from the literature . In α-Nb Si , the four non-equivalent sites
[41]
3
5
5
3
Nb , Nb , Si , and Si contain 588, 1,764, 784, and 392 data points, respectively. During the ML model
II
I
I
II
constructions in this work, 80% of the data were used for training and 20% for testing with random splits for
each dataset.
Features
The CE features, which encode information about local structure and composition, have been successfully
utilized to study alloys, oxides, and surface catalytic reactions [39,40] . The CE feature model can be given as an
(n + 1) dimensional compound feature vector as follows:
(1)
Here, P consists of n elementary features of an element or pure substance (P) and the target property T.
i
Each P is a two-dimensional vector representing the i-th elementary property, which includes the center
i
and environment components given by:
(2)
where
(3)
and
(4)
The normalized weight ω is defined as:
E,j
(5)

