Page 68 - Read Online
P. 68
Wen et al. J. Mater. Inf. 2025, 5, 30 https://dx.doi.org/10.20517/jmi.2024.102 Page 13 of 21
Figure 5. The true values and ML predicted values for (A) Hole reorganization energy, (B) Solvation free energy, (C) Maximum
absorption, and (D) LogP based on the RF model, respectively. ML: Machine learning; RF: random forest.
(0.996) > GBDT (0.992) > RF (0.991). It is evident that the XGBoost model outperforms all others in
predicting all four datasets. In terms of computational efficiency, the average training time for the three
models is as follows: XGBoost (8.2 s) < GBDT (206.3 s) < RF (611.3 s), highlighting the significantly higher
computational efficiency of the XGBoost model compared to GBDT and RF. Overall, the XGBoost model
demonstrates the best performance across the current dataset, with excellent generalization and
computational efficiency. The superior predictive accuracy and computational efficiency of XGBoost can be
attributed to its advanced algorithmic design. First, it employs an optimized gradient-boosted framework
integrated with L1/L2 regularization techniques to mitigate overfitting and enhance generalization
capabilities. Unlike RF, which relies on averaging multiple decision trees and may inadequately capture
complex feature interactions, XGBoost effectively models intricate relationships through its sequential tree-
building process. Second, it leverages a weighted quantile sketch for sparse data optimization and

