Page 70 - Read Online
P. 70
Wu et al. J. Mater. Inf. 2025, 5, 14 https://dx.doi.org/10.20517/jmi.2024.77 Page 7 of 15
Therefore, after considering both the predictive accuracy of MLFFs and the prediction speed of deep
learning models, we designed an integrated solution that uses MLFF as a data augmenter for the adsorption
energy prediction model. Initially, for low-index surfaces such as Cu(100), Cu(110), and Cu(111) depicted
in Figure 1B, where the enumerated configuration count is low, we randomly select 20% of the unoptimized
adsorption configurations for DFT calculations. For higher-index surfaces, we perform a primary sampling
based on adsorption site combination types, followed by a secondary sampling based on simplified
adsorption site combinations. The configurations obtained from secondary sampling undergo DFT
structural optimization. This sampling process yielded a total of 1,592 structures. The resulting structure-
energy pairs through structural optimization were obtained for training the MLFF [Supplementary Table 5].
Subsequently, we train the MLFF for the Cu-CO system with accuracy close to DFT calculations using the
open-source and user-friendly deep learning architecture DPMD, which has been proven to have high
simulation accuracy and efficiency in vast Si atomic systems [32,50] . This force field model is then applied to
optimize the remaining low-index surface configurations and those from the primary sampling. By
expanding the training dataset with adsorption configurations using MLFFs, we can significantly address
the issue of data scarcity faced when constructing deep learning models, thereby substantially improving
model quality .
[31]
Our trained force field exhibits high accuracy, with root mean square error (RMSE) values for energy and
force being < 1 meV/atom [Figure 3A] and < 0.03 eV/Å [Figure 3B], respectively. The optimization of initial
configuration guesses in the secondary sampling space using the well-fitted force field further assesses the
force field quality: the steady-state adsorption configurations obtained from force field optimization almost
perfectly match those obtained from DFT optimization [Figure 3C]. Moreover, analyzing the distribution of
Cu–C bond lengths in the steady-state adsorption configuration set reveals that the distribution from MLFF
optimization closely aligns with that from DFT optimization [Figure 3D], with an average Cu–C bond
length discrepancy of < 0.01 Å. This demonstrates that the MLFF can achieve near-DFT calculation
accuracy for adsorption configurations across eight Cu surfaces. The high-quality force field is attributed
not only to the superior MLFF architecture but also to the effective sampling method - ensuring the
sampled configuration guesses, i.e., the starting points for DFT geometric optimization, are as diverse as
possible, facilitating a rich optimization trajectory that benefits model learning of more complex multi-body
interactions. We used the MFLL to optimize ~186,000 structures for the following graph neural network
(GNN) model and the distribution of structures was presented in Supplementary Table 6. The specific
training parameters of MLFF are shown in Supplementary Table 7.
Graph-based model for adsorption energy prediction
In recent years, GNN models have shone brightly in material design and performance prediction, leveraging
graph embeddings of materials for various predictive tasks [51-54] . One of their key advantages over other deep
learning models lies in the natural suitability of graph structures for describing chemical and material
structures, with nodes and edges representing atoms and their interactions, respectively. Hence, extracting
high-quality graph data from materials is crucial. In our study, the interactions among CO molecules affect
their adsorption energy on Cu surfaces, especially at high coverages. Previous research often overlooked
non-bonding interactions beyond hydrogen bonds when constructing graph data, failing to accurately
[18]
characterize adsorption structures . We propose an adsorption configuration graph data extraction
method that effectively captures the impact of CO interactions on their adsorption energy, as shown in
Figure 4A and Supplementary Figure 2. The first step involves searching for atoms in contact with a CO
molecule on the Cu surface, considering it as the central CO, within the van der Waals radius to form a
neighbor atom set, including first-order neighbor COs in contact. The second step continues the search for
atoms in contact with these first-order neighbor COs and builds local subgraphs for multiple neighbor sets
centered on CO. The third step merges the central CO and its neighboring COs’ local subgraphs into a

