Page 37 - Read Online
P. 37
Page 10 of 16 Lu et al. J Mater Inf 2024;4:31 https://dx.doi.org/10.20517/jmi.2024.65
Although the computational hydrogen electrode (CHE) model can effectively describe the reaction
mechanism and catalytic activity of NRR on Nb@C N-NCM, it is overly simplistic for complex working
2
conditions, considering the influence of electrode potential only through energy correction. Therefore, we
further employed the Standard Hydrogen Electrode (SHE) model to investigate the effect of potential on the
NRR activity of Nb@C N-NCM. Supplementary Figure 3A shows the computed energies as a function of
2
the applied electrode potential (vs. SHE) for Nb@C N-NCM and the corresponding reaction intermediates.
2
It demonstrates that the energy-potential points fit well into a quadratic function. As shown in
Supplementary Figure 3B, we obtained the electrode potential-dependent free energy curves. The results
indicate that Nb@C N-NCM exhibits the best catalytic activity at an electrode potential of -3V, with the
2
corresponding U determined to be -0.14 V vs. SHE. Subsequently, we considered the effect of potential on
L
the selectivity of the catalyst [Supplementary Figure 3C]. The results show that within the potential range
considered, N always has lower E and more readily occupies the active sites of the catalyst compared to
2
ads
the H atom, indicating excellent selectivity for NRR over HER.
Machine learning analysis
To explore the intrinsic factors influencing the catalytic performance of NRR catalysts, ML was employed to
uncover the relationships between fundamental physicochemical properties and catalytic activity. The
workflow of the ML approach is shown in Figure 7. The target variable dataset was derived from the
DFT-calculated results. Based on previous studies and validation in this work, the first and last protonation
steps were identified as key steps for assessing catalyst activity. These steps are collectively referred to as
Candidate Potential Determining Steps (C-PDS), and were used as the target variable for ML analysis. The
final dataset includes 28 ΔG values for the end-on mode first protonation, 24 for the side-on mode first
protonation, and 28 for the final protonation step.
To construct a reliable feature set, 20 features were selected (see Supplementary Table 2). Seventeen of these
features represent the inherent properties of TM atoms, such as Pauling electronegativity (χ ), the number of
P
d-electrons (N ), and electron affinity (EA), with values obtained from the National Institute of Standards
d
[78]
and Technology (NIST) database .
Additionally, three binary features were introduced to distinguish between different protonation steps and
adsorption modes: S (first-step protonation of side-on adsorption), S (first-step protonation of end-on
1S
1E
adsorption), and S (last-step protonation). These features were encoded as one-hot vectors and
6
incorporated into the feature set. We first analyzed the Pearson correlation among the 20 features to identify
any potential redundancy. Subsequently, the RFE method was employed to select the optimal feature subset.
The Pearson correlation heatmap is presented in Supplementary Figure 4, while the final dataset after
feature selection is summarized in Supplementary Table 3. The results of optimal ML models using three
algorithms (XGBR, GBR, RFR) are shown in Figure 8A-C. All models demonstrated strong linear
2
correlations between the predicted values and the DFT-calculated results, with R values ranging from 0.88
to 0.91 and MAE values between 0.19 and 0.24, indicating excellent predictive performance across all
2
models. Among these, XGBR exhibited the best performance, achieving the highest R and the lowest MAE.
Additionally, the average R and MAE values from the 5-fold cross-validation are summarized in Figure 8D,
2
further confirming XGBR as the best-performing model. Furthermore, we conducted additional model
training using eight features for comparison. The results, as shown in Supplementary Figure 5, indicate that
the average scores are very close, thereby supporting the reliability of the model.
Feature importance analysis revealed that the one-hot encoded features S contributed the most, accounting
6
for 56.96% of the total importance. These features are crucial for distinguishing between the first and last

