Page 115 - Read Online
P. 115
Page 8 of 15 Shang et al. J. Mater. Inf. 2025, 5, 52 https://dx.doi.org/10.20517/jmi.2025.36
Figure 3. Performance comparison of different ML models. (A) MAE bar chart for three model types: RF, GBDT, and LGB; (B) ROI bar
chart for three model types: RF, GBDT, and LGB. ML: Machine learning; MAE: mean absolute error; RF: random forest; GBDT: gradient
boosting decision tree; LGB: lightGBM; ROI: relative overfitting index.
Figure 4. Feature importance analysis using SHAP values. (A) SHAP value plot showing the impact of various features on the model
output; (B) Bar chart of mean SHAP values for different features, with a pie chart illustrating the proportion of descriptor categories:
SISSO descriptors (S), functional group descriptors (T), and other descriptors (O). SHAP: SHapley Additive exPlanations; SISSO: Sure
Independence Screening and Sparsifying Operator.
values and observing the impact on model performance. In Figure 4B, descriptors-S refers to the SISSO
descriptor, descriptors-T encompasses descriptors associated with functional groups, and descriptors-O
comprises additional descriptor types. The feature importance analysis demonstrates that SISSO descriptors
hold the top three positions, contributing a significant 61.6% to the model’s performance, thereby acting as a
crucial component for enhancing overall model performance. It is worth noting that three out of the four
most important features immediately following the SISSO descriptors are associated with the functional
groups of MXenes. Moreover, the three selected SISSO descriptors that exhibit the highest correlation with
the work function were also centered around the properties of these functional groups. Consequently, the
SISSO approach proves effective in identifying key features that not only enhance the reliability and stability
of the model, but also reveal a strong correlation between the work function of MXenes and their surface
functional groups.
To validate the previous conclusion that SISSO descriptors and functional group descriptors consistently
play a crucial role across all models, we verified their interpretability by comparing them with other models.

