Page 253 - Read Online
P. 253

Page 18 of 29                                                      Liu et al. J. Mater. Inf. 2026, 6, 18









































               Figure 10. A performance comparison chart of 12 machine learning model test sets constructed based on the optimal plastic strain feature
               subset. ML: Machine learning; R : the coefficient of determination; RMSE: root mean square error; RF: random forest; Bagging: bootstrap
                                    2
               aggregating; ExtraTree: extremely randomized tree; XGBoost: eXtreme gradient boosting; DT: decision tree; KNN: K-nearest neighbor;
               GBRT: gradient boosting regression tree; LGBM: light gradient boosting machine; AdaBoost: adaptive boosting; ANN: artificial neural
               network; SVR: support vector regression.

               Table 6. Hyperparameter search ranges for the RF model

               Name              Parameter description                                   Min   Max   Step
               n_estimators      The number of DTs in a RF                               50    550   50
               max_depth         The maximum depth of the DT                             0     30    5
               min_samples_split  The minimum number of samples required to split internal nodes  2  20  2
               min_samples_leaf  The minimum number of samples required for leaf nodes   1     10    1

               RF: Random forest; DTs: decision trees.

               ductility (RF) models were analyzed [Figure 12]. Both models exhibited rapid convergence, with test set
               errors stabilizing after approximately 50 iterations. Notably, despite the typical gap between training and
               testing errors inherent in ensemble methods, the test error curves remained flat without upward trends as
               model complexity increased. This behavior confirms that the models effectively learned the underlying
               composition-property relationships without overfitting to noise, ensuring their generalizability to unseen
               data.


               ML model construction for multi-objective optimization and SHapley Additive exPlanations
               importance analysis
               To construct predictive models based on composition-domain knowledge features, proportional elemental
               contents are integrated with domain-derived key features to form composite descriptors. This integration not
               only satisfies the intrinsic requirements of genetic algorithms for multi-objective optimization, but also
               incorporates domain knowledge to enhance model accuracy and interpretability. To better fit the
   248   249   250   251   252   253   254   255   256   257   258