Page 68 - Read Online
P. 68

Wen et al. J. Mater. Inf. 2025, 5, 30  https://dx.doi.org/10.20517/jmi.2024.102  Page 13 of 21


























































                Figure 5. The true values and ML predicted values for (A) Hole reorganization energy, (B) Solvation free energy, (C) Maximum
                absorption, and (D) LogP based on the RF model, respectively. ML: Machine learning; RF: random forest.

               (0.996) > GBDT (0.992) > RF (0.991). It is evident that the XGBoost model outperforms all others in
               predicting all four datasets. In terms of computational efficiency, the average training time for the three
               models is as follows: XGBoost (8.2 s) < GBDT (206.3 s) < RF (611.3 s), highlighting the significantly higher
               computational efficiency of the XGBoost model compared to GBDT and RF. Overall, the XGBoost model
               demonstrates the best performance across the current dataset, with excellent generalization and
               computational efficiency. The superior predictive accuracy and computational efficiency of XGBoost can be
               attributed to its advanced algorithmic design. First, it employs an optimized gradient-boosted framework
               integrated with L1/L2 regularization techniques to mitigate overfitting and enhance generalization
               capabilities. Unlike RF, which relies on averaging multiple decision trees and may inadequately capture
               complex feature interactions, XGBoost effectively models intricate relationships through its sequential tree-
               building process. Second, it leverages a weighted quantile sketch for sparse data optimization and
   63   64   65   66   67   68   69   70   71   72   73