Page 71 - Read Online
P. 71
Page 8 of 15 Wu et al. J. Mater. Inf. 2025, 5, 14 https://dx.doi.org/10.20517/jmi.2024.77
Figure 3. Evaluation of MLFF accuracy. (A) Scatter plot of the energy predictions, demonstrating the MLFF’s high accuracy, with an
RMSE for energy within 1 meV/atom, as validated against DFT calculations; (B) Corresponding force predictions plot, where the MLFF
achieves an RMSE of less than 0.03 eV/Å, indicating the model’s precision in capturing the forces in the Cu-CO system across a vast
configuration space; (C) Comparing the initial CO adsorption configuration with those optimized by DFT and MLFF, showing the MLFF’s
ability to replicate DFT-level structural accuracy on Cu surfaces; (D) Histogram of Cu–C bond lengths from stable-state configurations,
highlighting the close match between MLFF and DFT results, with bond length discrepancies averaging less than 0.01 Å. MLFF: Machine-
learning force field; DFT: density functional theory; RMSE: root mean square error.
feature graph representing the adsorption environment of the central CO. Different interaction types in the
feature graph are assigned distinct indicator vectors during encoding, including non-bonding interactions
between CO molecules, akin to the encoding of hydrogen bonds in previous research .
[26]
Despite using only a few very simple attribute features [Supplementary Table 8] and conventional training
steps [Supplementary Table 9], the GNN model demonstrated excellent predictive accuracy and
generalization ability across eight index surfaces at varying coverages [Supplementary Figure 3]. The
model’s superior performance stems partly from more accurate graph descriptors. Unlike models that
ignored CO non-bonding interactions when constructing graph data, our model significantly improved
performance [Supplementary Figure 4]. Additionally, it benefited from the richness of training graph data,
thanks to the expansion of training data via MLFF, presenting an advantage over models trained solely on
data from DFT calculations. The accuracy obtained by training solely on DFT data [Figure 4B] is
significantly lower than that achieved through training on an expanded dataset (about 186,000) under the
same model architecture [Figure 4C]. The latter exhibits improvements in accuracy by 16% and 28% in
terms of the coefficient of determination (R ) and root mean squared error (RMSE), respectively
2

