Page 253 - Read Online
P. 253
Page 18 of 29 Liu et al. J. Mater. Inf. 2026, 6, 18
Figure 10. A performance comparison chart of 12 machine learning model test sets constructed based on the optimal plastic strain feature
subset. ML: Machine learning; R : the coefficient of determination; RMSE: root mean square error; RF: random forest; Bagging: bootstrap
2
aggregating; ExtraTree: extremely randomized tree; XGBoost: eXtreme gradient boosting; DT: decision tree; KNN: K-nearest neighbor;
GBRT: gradient boosting regression tree; LGBM: light gradient boosting machine; AdaBoost: adaptive boosting; ANN: artificial neural
network; SVR: support vector regression.
Table 6. Hyperparameter search ranges for the RF model
Name Parameter description Min Max Step
n_estimators The number of DTs in a RF 50 550 50
max_depth The maximum depth of the DT 0 30 5
min_samples_split The minimum number of samples required to split internal nodes 2 20 2
min_samples_leaf The minimum number of samples required for leaf nodes 1 10 1
RF: Random forest; DTs: decision trees.
ductility (RF) models were analyzed [Figure 12]. Both models exhibited rapid convergence, with test set
errors stabilizing after approximately 50 iterations. Notably, despite the typical gap between training and
testing errors inherent in ensemble methods, the test error curves remained flat without upward trends as
model complexity increased. This behavior confirms that the models effectively learned the underlying
composition-property relationships without overfitting to noise, ensuring their generalizability to unseen
data.
ML model construction for multi-objective optimization and SHapley Additive exPlanations
importance analysis
To construct predictive models based on composition-domain knowledge features, proportional elemental
contents are integrated with domain-derived key features to form composite descriptors. This integration not
only satisfies the intrinsic requirements of genetic algorithms for multi-objective optimization, but also
incorporates domain knowledge to enhance model accuracy and interpretability. To better fit the

