Page 109 - Read Online
P. 109
Page 2 of 15 Shang et al. J. Mater. Inf. 2025, 5, 52 https://dx.doi.org/10.20517/jmi.2025.36
interpretability in ML predictions of MXenes’ work functions, providing both fundamental insights and practical
tools for materials discovery.
Keywords: Machine learning, MXenes, interpretability, SHAP, SISSO, work function
INTRODUCTION
Two-dimensional (2D) carbides and nitrides, MXenes, are an emerging class of low-dimensional materials
[1,2]
that have received considerable interest over the last decade , which exhibit tunable compositions and rich
surface chemistry . The tunability of MXenes allows for the optimization of their electrical conductivity,
[3-5]
mechanical robustness, and chemical stability, making them indispensable in widespread applications . In
[6-8]
particular, their work function can be tuned over a wide range from 1.3 to 7.2 eV, making MXenes highly
promising candidates in the fields of optoelectronics, catalysis, sensing, etc. [9-12] . In photonic devices, the
work function is a critical parameter that influences the device’s performance. However, performing density
functional theory (DFT) calculations for each candidate material is impractical because of the enormous
computational and time costs, which renders traditional trial-and-error approaches less feasible.
The rapid advancements of machine learning (ML) have established it as a transformative tool for
accelerating material science research [13-16] . Leveraging large-scale datasets, ML has emerged as a crucial
enabler in predicting material properties and uncovering insights into complex systems. For example, ML
[17]
models have been used not only to evaluate the stability of 2D materials , but also to facilitate swift
assessments of novel compounds with potential applications in photocatalysis and energy storage
devices . Expanding upon these capabilities, AI-driven frameworks have successfully addressed complex
[18]
challenges in material design, and achieved significant achievements. For example, the development of ML
models, including gradient boosting and symbolic regression, with DFT to accelerate the design of MXene
[19]
and MN4-graphene bifunctional oxygen electrocatalysts , the creation of the CrabNet neural network
architecture to predict material properties from compositions, enhancing regression performance through
attention mechanisms and element embeddings , and the discovery of stable AA0MH6 semiconductors
[20]
via generative adversarial networks (GAN) coupled with DFT validations, etc. . Therefore, ML can provide
[21]
an efficient strategy for screening and predicting materials with desired work functions.
However, although Roy et al. have employed ML methods to predict the work function of MXenes and
[22]
reduced the mean absolute error (MAE) to approximately 0.26 eV , there are still some limitations. First,
further improving the predictive accuracy of the models remains a challenge, limited by the diversity of
available algorithms and the size of valid data . Secondly, further enhancing the interpretability of ML
[23]
models remains a challenge, limited by their “black box” character, making it difficult to understand the
intrinsic relationship between material features and target performance . Therefore, there is an urgent
[24]
need to develop a ML model integrated with advanced algorithms that not only enhances the prediction
accuracy of the work function of MXenes but also improves the model’s interpretability, thereby elucidating
the intrinsic connection between material characteristics and the work function.
In this study, we integrate the Sure Independence Screening and Sparsifying Operator (SISSO) method with
the stacked model approach to improve the prediction accuracy of the work function of MXenes. The SISSO
method generates highly correlated descriptors of the target features, which provide a transparent and
intuitive overview of the relationship between the target features and the physical properties. The stacked
model further improves the model performance by integrating the prediction results from multiple base
models and using them as inputs for secondary learning. By interpreting the ML model using the SHapley

