Page 66 - Read Online
P. 66
Wu et al. J. Mater. Inf. 2025, 5, 14 https://dx.doi.org/10.20517/jmi.2024.77 Page 3 of 15
method, we selected a very small set of configurations for DFT structural optimization to obtain
optimization trajectories for training a MLFF based on deep potential molecular dynamics (DPMD).
Subsequently, the MLFF was used to perform structural optimizations on an expanded sampling space
(~186,000) to obtain corresponding stable adsorption energies. Finally, we trained a graph embedding
network model (GEN) using the configuration-energy data, which considered non-bonded adsorbate
interactions in feature construction and efficiently predicted energies across the entire target configuration
space. Applied to the Cu-CO system, our method achieved results qualitatively consistent with experiments
at three orders of magnitude lower computational cost than pure DFT calculations: the adsorption strength
of CO on Cu surfaces exhibited a minimal change in energy at first, followed by a significant increase with
coverage; high-index Cu surfaces often exhibited higher catalytic activity due to more low-coordination Cu
sites. These findings undoubtedly demonstrate the effectiveness and efficiency of our method and its power
in exploring vast configuration spaces.
MATERIALS AND METHODS
At high coverage, the configurations of adsorption not only experience an explosive increase in number due
to the enumerated surfaces and the geometric structures and binding modes of the adsorbates but also
become almost unpredictable due to the complex interactions between the adsorbates. This signifies that the
[25]
enumeration methods relying solely on expert experience fail under these circumstances . By integrating
expert knowledge with deep learning technologies, a programmable scalable agent model can provide
interpretable and reliable analysis and predictions for the vast configuration space.
Workflow
Our study focuses on the adsorption configurations of CO on eight different Cu surfaces at varying
coverage levels. As illustrated in Figure 1A, taking the (100) surface as an example, as the CO coverage
increases, the distance between the adsorbed CO molecules gradually decreases, along with an increase in
the interaction between the adsorbates. The number of configurations shows a trend of initially increasing
and then decreasing [Supplementary Table 1]. Ultimately, nearly 7 million adsorption configurations are
generated for the eight Cu surfaces, representing an extremely large configuration space [Supplementary
Table 2]. We will introduce the detailed enumeration process in the following sections.
We propose a simple and efficient framework capable of rapidly and accurately predicting the CO
adsorption energies of all stable configurations. The workflow is illustrated in Figure 1B, which depicts our
comprehensive computational exploration of the configuration space. Initially, the entire configuration
space is sampled randomly twice to obtain the first and second sampling spaces, respectively, with both
sampling steps covering all surface coverage levels. Subsequently, configurations from the more concise
second sampling space undergo DFT structural optimization to obtain the CO adsorption energies of stable
configurations, along with the trajectory of configuration optimization and the corresponding energies.
Using these trajectories and energies as a dataset, a MLFF based on the DPMD deep potential architecture is
trained, which effectively fits the interaction between adsorbates at different coverages . The fitted force
[32]
field is then used to optimize the structures of the larger first sampling space to obtain the CO adsorption
energies of stable configurations. The resulting configuration-adsorption energy data can train a well-
performing GEN, capable of accurately predicting the stable CO energies of approximately 7 million
adsorption configurations in the target configuration space, despite using very simple feature combinations.
Compared to the total configurational space, the computational framework requires a significantly lower
volume of initial data from DFT calculations, differing by three orders of magnitude, even though DFT
calculations are generally considered to be highly resource-intensive. By introducing a machine-learned

