Page 55 - Read Online
P. 55
Page 12 of 20 Zhu et al. J. Mater. Inf. 2025, 5, 8 https://dx.doi.org/10.20517/jmi.2024.76
experimental density values and those predicted by the most effective machine learning model for
intermetallic compounds. Figure 3A shows the machine learning model predictions for the Hexagonal
system. The bar chart indicates that among the five machine learning models, XGBoost achieves the highest
2
2
accuracy with R = 0.9798, while the SVM model has the lowest accuracy with R = 0.8538. The scatter plot
in Figure 3A compares XGBoost predictions with calculated values by DFT, where the points cluster closely
around the diagonal line (y = x), demonstrating a strong agreement between the predicted and experimental
densities. XGBoost still shows the best prediction accuracy for the cubic system, with R close to 0.98, as
2
shown in Figure 3B. However, for the monoclinic system, R = 0.9825 is achieved by the RF model, which
2
outperforms XGBoost, as depicted in Figure 3C. Figure 3D-F shows that XGBoost achieves the highest
2
accuracy for the remaining systems including orthorhombic, trigonal, and tetragonal systems with R values
of 0.9692, 0.9825, and 0.9826, respectively. From Figure 3, it can be seen that the data points for each model
(especially XGBoost) are mostly aligned along the diagonal line, indicating that the predictions are generally
accurate and demonstrating the potential of machine learning models for predicting the density of
intermetallic compounds.
Machine learning modeling on aggregated crystal structures
In the above section, machine learning models are built based on data subsets for specific crystal structures,
achieving high prediction accuracy. However, this single crystal structure-based modeling approach has
limitations in practical applications, especially when studies require density prediction on mixed data of
multiple crystal structures. In this section, an approach without distinguishing crystal structures is adopted,
integrating the data from all seven crystal structures for unified modeling to explore the performance of
traditional machine learning models on this complex dataset. The model results are shown in Figure 4A and
B. Among the traditional machine learning models, the accuracy of simple models such as LR or KNN
remained almost the same or slightly increased, while the prediction accuracy of complex models such as RF
or XGBoost decreased to varying degrees. For example, the accuracy of the XGBoost model that shows
optimal performance in the above section decreased from 0.98 to 0.9468. Figure 4C-E shows scatter plots of
the predictions from the LR, SVM, and XGBoost models. From these images, it can be seen that the
prediction results of the LR and SVM models are poor, with some points deviating far from the diagonal
line. The XGBoost model has higher accuracy than the other two, but a few data points deviate significantly
from the diagonal, reducing prediction accuracy. A review of these compounds revealed that they all have
polymorphic forms in the dataset. For instance, the test set includes the compound In Hg with an
3
3
“Orthorhombic” structure and an experimental density of 1.4236 g/cm . However, in the training set, In Hg
3
is classified under a “Hexagonal” structure with a significantly higher experimental density of 8.6732 g/cm .
3
This discrepancy likely contributes to the XGBoost model inaccurately predicting the density of In Hg as
3
7.5879 g/cm in the test set. The compound PuPt , which exhibits a “Tetragonal” structure and a density of
3
3
3
18.6427 g/cm in the training set, appears with a “Cubic” structure and a density of 19.9649 g/cm in the test
3
set. The XGBoost prediction is 17.1343 g/cm , aligning more closely with the training set density.
3
The above analysis shows that intermetallic density data for different crystal structures vary significantly.
Directly combining all structures’ data for modeling causes traditional models to struggle in recognizing
differences between crystal structures, which in turn affects prediction performance. In this context, The
study introduces an IGNN model that is better suited for multi-crystal structure data, addressing the
limitations of traditional models when applied to mixed crystal structure datasets. As illustrated in
Figure 4A and B, the constructed IGNN model clearly outperforms traditional machine learning models,
achieving the highest R value of 0.9884 and the lowest RMSE of 0.4291 g/cm . Additionally, the scatter plot
2
3
in Figure 4F shows that the IGNN model can utilize node information in the graph structure to
automatically recognize differences between crystal structures and incorporate this information as input for
modeling, enhancing its adaptability to multi-structure datasets. The prediction error for the compound

