Page 247 - Read Online
P. 247

Page 12 of 29                                                      Liu et al. J. Mater. Inf. 2026, 6, 18























               Figure 4. Preliminary screening results of strength feature engineering. (A) Pearson correlation coefficient matrix among features; (B) MI
               between features and target variables; (C) Feature importance ranking based on the RF model. MI: Mutual information; RF: random forest;
               EWF: electron work function; B/G: bulk/shear modulus; VEC: valence electron concentration.


               retaining the five most significant parameters {Strain rate, EWF, Fermi, ΔHmix, Ω}, which collectively
               account for over 90% of the total importance. All elemental features are removed after Pearson correlation
               coefficient screening, indicating that their predictive contribution is significantly lower than that of
               knowledge-based physical descriptors. These results suggest that physics-informed descriptors provide
               higher prediction accuracy. As shown in Figure 5, an exhaustive subset evaluation using a wrapper method is
               conducted to identify the key performance parameters (KPPs) influencing titanium alloy strength. The
               optimal three-feature combination {Strain rate, Fermi, Ω} is determined, achieving a linear correlation
               coefficient of 0.89. This configuration balances model complexity and predictive accuracy, with an R  of
                                                                                                        2
               0.886, representing only a marginal decrease (0.1%) compared to the four-feature subset {Strain rate, Fermi,
               ΔHmix, Ω} with R  = 0.887, while significantly reducing the risk of overfitting. Subsets incorporating strain
                              2
               rate demonstrated superior performance in fitting the training data compared to those excluding it [Table 4],
               indicating more effective learning of deformation patterns. Models without strain rate exhibit overfitting,
               with substantial discrepancies between training and testing errors. For example, the subset {EWF, Fermi, Ω}
               achieved a training set accuracy of 0.78 and a testing set accuracy of 0.88, emphasizing the critical importance
               of strain rate as an essential feature in predictive models.

               For the ductility dataset, initial feature selection was conducted using Pearson correlation coefficients, MI,
               and RF feature importance ranking [Figure 6A and B]. A subset comprising ten parameters {Strain rate,
               B/G, Zr, ΔHmix, Fe, V, VEC, Sn, Nb, Si} is identified [Figure 6C]. Subsequently, six features {Strain rate,
               B/G, Zr, ΔHmix, Fe, V}, collectively accounting for over 90% of the cumulative importance, are retained
               through feature importance analysis, indicating their dominant influence on ductility. Among these, strain
               rate alone shows an importance greater than 0.55 under the experimental conditions. To reduce
               computational complexity during feature subset evaluation, these six features were further subjected to
               exhaustive subset screening[Figure 7 and Table 5]. Single-feature subsets achieve limited predictive accuracy
               (R  = 0.56), whereas multi-feature configurations maintain higher predictive performance (R  = 0.82). Subsets
                                                                                            2
                 2
               including strain rate consistently exhibit training R  values above 0.85. Three- and four-parameter subsets
                                                           2
               containing strain rate demonstrate a clear trend of strong predictive performance (R  > 0.85), leading to the
                                                                                       2
               identification of five optimal subsets: {Strain rate, B/G, ΔHmix}, {V, Strain rate, ΔHmix}, {V, Fe, Strain rate,
               ΔHmix}, {V, Strain rate, B/G, ΔHmix}, and {Fe, Strain rate, B/G, ΔHmix}. Among these, the three-feature
               subset {Strain rate, B/G, ΔHmix} provides the best balance between predictive accuracy and model
               interpretability.


               Proactive exploration of optimal ML models and targeted optimization
               Hyperparameter optimization directly influences the performance ceiling and generalization capability of
               ML algorithms, making it critical for achieving optimal task-specific performance.
   242   243   244   245   246   247   248   249   250   251   252