Page 247 - Read Online
P. 247
Page 12 of 29 Liu et al. J. Mater. Inf. 2026, 6, 18
Figure 4. Preliminary screening results of strength feature engineering. (A) Pearson correlation coefficient matrix among features; (B) MI
between features and target variables; (C) Feature importance ranking based on the RF model. MI: Mutual information; RF: random forest;
EWF: electron work function; B/G: bulk/shear modulus; VEC: valence electron concentration.
retaining the five most significant parameters {Strain rate, EWF, Fermi, ΔHmix, Ω}, which collectively
account for over 90% of the total importance. All elemental features are removed after Pearson correlation
coefficient screening, indicating that their predictive contribution is significantly lower than that of
knowledge-based physical descriptors. These results suggest that physics-informed descriptors provide
higher prediction accuracy. As shown in Figure 5, an exhaustive subset evaluation using a wrapper method is
conducted to identify the key performance parameters (KPPs) influencing titanium alloy strength. The
optimal three-feature combination {Strain rate, Fermi, Ω} is determined, achieving a linear correlation
coefficient of 0.89. This configuration balances model complexity and predictive accuracy, with an R of
2
0.886, representing only a marginal decrease (0.1%) compared to the four-feature subset {Strain rate, Fermi,
ΔHmix, Ω} with R = 0.887, while significantly reducing the risk of overfitting. Subsets incorporating strain
2
rate demonstrated superior performance in fitting the training data compared to those excluding it [Table 4],
indicating more effective learning of deformation patterns. Models without strain rate exhibit overfitting,
with substantial discrepancies between training and testing errors. For example, the subset {EWF, Fermi, Ω}
achieved a training set accuracy of 0.78 and a testing set accuracy of 0.88, emphasizing the critical importance
of strain rate as an essential feature in predictive models.
For the ductility dataset, initial feature selection was conducted using Pearson correlation coefficients, MI,
and RF feature importance ranking [Figure 6A and B]. A subset comprising ten parameters {Strain rate,
B/G, Zr, ΔHmix, Fe, V, VEC, Sn, Nb, Si} is identified [Figure 6C]. Subsequently, six features {Strain rate,
B/G, Zr, ΔHmix, Fe, V}, collectively accounting for over 90% of the cumulative importance, are retained
through feature importance analysis, indicating their dominant influence on ductility. Among these, strain
rate alone shows an importance greater than 0.55 under the experimental conditions. To reduce
computational complexity during feature subset evaluation, these six features were further subjected to
exhaustive subset screening[Figure 7 and Table 5]. Single-feature subsets achieve limited predictive accuracy
(R = 0.56), whereas multi-feature configurations maintain higher predictive performance (R = 0.82). Subsets
2
2
including strain rate consistently exhibit training R values above 0.85. Three- and four-parameter subsets
2
containing strain rate demonstrate a clear trend of strong predictive performance (R > 0.85), leading to the
2
identification of five optimal subsets: {Strain rate, B/G, ΔHmix}, {V, Strain rate, ΔHmix}, {V, Fe, Strain rate,
ΔHmix}, {V, Strain rate, B/G, ΔHmix}, and {Fe, Strain rate, B/G, ΔHmix}. Among these, the three-feature
subset {Strain rate, B/G, ΔHmix} provides the best balance between predictive accuracy and model
interpretability.
Proactive exploration of optimal ML models and targeted optimization
Hyperparameter optimization directly influences the performance ceiling and generalization capability of
ML algorithms, making it critical for achieving optimal task-specific performance.

